EDBT 2026 Demo / reviewers in the wild / expert
Yunxin Liu 0001
dblp:55/3521-1
· DBLP profile ↗
140ranked-venue papers
2as first author
81since 2021 · last 2026
0000-0001-7352-8955ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 84 · 2 first-author · 44 since 2021Systems, architecture and hardware · 19 · 16 since 2021Software engineering, systems software and programming languages · 15 · 9 since 2021Artificial intelligence and machine learning · 9 · 8 since 2021Databases, data management, data science and information retrieval · 6 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 since 2021Security and privacy · 4 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 since 2021Human-computer interaction and ubiquitous computing · 3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enhancing Agentic Textual Graph Retrieval with Synthetic Stepwise SupervisionabstractIntegrating textual graphs into Large Language Models (LLMs) is promising for complex graph-based QA. However, a key bottleneck is retrieving informative yet compact subgraphs that fit the LLM context. Existing retrievers often struggle, relying either on shallow embedding similarity or costly interactive policies that require excessive supervision. To address these challenges, we introduce an agentic textual graph reasoning framework featuring an LLM-based retriever trained with synthetic stepwise supervision. Rather than relying on final answer rewards which often yield sparse and unstable signals, we optimize the retriever by evaluating each step against offline-extracted golden subgraphs. Our approach distills golden subgraphs via a specialized data synthesis pipeline to formulate dense rewards, facilitating a two-stage training scheme that effectively learns the interactive graph exploration policy. Based on extensive experiments on three common datasets in comparison with seven strong baselines, our approach achieves an average improvement of 15.6% in accuracy and 17.2% in F1 score. The advantage is even higher in more complicated multi-hop reasoning tasks. Ge Chang 0002, Jinbo Su, Yuhao Shang, Huiwen Zheng, Hongli Ma, Yuanchun Li 0003, Yunxin Liu 0001 |
ACL (1) | 10 |
| 2026 | Benchmarking LLM's Capability in Reasoning over Conflicting Web ReferencesabstractLarge language models (LLMs) integrated with retrieval-augmented generation (RAG) have become a dominant framework for building intelligent assistants.In real-world applications such as ChatGPT with web search, the retrieved document often comes from diverse, potentially unreliable sources and may contain inconsistent claims.Unlike traditional search engines that rely on users to manually compare information, LLM-based systems typically feed all retrieved content into the model's context, requiring LLMs to autonomously identify, differentiate, and reason over conflicting viewpoints.Unlike mainstream LLM evaluation tasks like math and code generation that are primarily focused on reasoning with factual context, question-answering with multi-source references requires fundamentally different capabilities to identify and reason over knowledge contradictions.In this paper, we introduce CONFRAG, a benchmark for evaluating LLMs' reasoning capability over real-world conflicting documents retrieved from the web.It consists of 1,814 real-world questions, each paired with an average of 9.58 retrieved paragraphs from heterogeneous online sources.A total of 57.2% of the questions exhibit explicit contradictions.We further propose three structured evaluation tasks, answer clustering, answer coverage, and reason coverage, to quantify a model's ability to organize and explain contradictory content.Experiments with state-of-the-art models such as GPT-4.1 and Claude-3-7-Sonnet reveal substantial performance gaps, highlighting the need for more targeted research in contradiction-aware question answering.To the best of our knowledge, CONFRAG is the first benchmark specifically designed to evaluate contradiction-aware reasoning on real-world long web documents. Yizhen Yuan, Yuanchun Li 0003, Yunxin Liu 0001 |
ACL (1) | 5 |
| 2026 | VoLLM: Smoothness-aware Serving of LLM-powered Voice Q&A via Adaptive Preemption
Yuanchun Li 0003, Ju Ren 0001, Lichen Pang, Shansong Yang, Yunxin Liu 0001 |
IPDPS | 6 |
| 2026 | Vec-LUT: Vector Table Lookup for Parallel Ultra-Low-Bit LLM Inference on Edge DevicesabstractLarge language models (LLMs) are increasingly deployed on edge devices. To meet strict resource constraints, real-world deployment has pushed LLM quantization from 8-bit to 4-bit, 2-bit, and now 1.58-bit. Combined with lookup table (LUT)-based inference, CPUs run these ultra-low-bit LLMs even faster than NPUs, opening new opportunities for ubiquitous on-device intelligence. Weijun Wang 0001, Jianyu Wei, Ting Cao 0003, Yunxin Liu 0001 |
MobiSys | 6 |
| 2026 | Mobile GUI Agents under Real-world Threats: Are We There Yet?abstractRecent years have witnessed a rapid development of mobile GUI agents powered by large language models (LLMs), which can autonomously execute diverse device-control tasks based on natural language instructions. The increasing accuracy of these agents on standard benchmarks has raised expectations for large-scale real-world deployment, and there are already several commercial agents released and used by early adopters. However, are we really ready for GUI agents integrated into our daily devices as system building blocks? We argue that an important pre-deployment validation is missing to examine whether the agents can maintain their performance under real-world threats. Specifically, unlike existing common benchmarks that are based on simple static app contents (they have to do so to ensure environment consistency between different tests), real-world apps are filled with contents from untrustworthy third parties, such as advertisement emails, user-generated posts and medias, etc. These contents may inevitably appear in the agents' observation space and influence the task execution process. Systematic investigation of this problem is challenging since the real-world app contents are significantly skewed—testing on normal real-world apps usually cannot uncover any potential risk since most app contents are benign. To this end, we introduce a scalable app content instrumentation framework to enable flexible and targeted content modifications within existing applications. Leveraging this framework, we create a test suite comprising both a dynamic task execution environment and a static dataset of challenging GUI states. The dynamic environment encompasses 122 reproducible tasks, and the static dataset consists of over 3,000 scenarios constructed from commercial apps. We perform experiments on both open-source and commercial GUI agents. Our findings reveal that all examined agents can be significantly degraded due to third-party contents, with an average misleading rate of 42.0% and 36.1% in dynamic and static environments respectively. The framework and benchmark has been released at https://agenthazard.github.io. Guohong Liu 0002, Jialei Ye, Wei Liu 0302, Pengzhi Gao, Jian Luan 0001, Yuanchun Li 0003, Yunxin Liu 0001 |
MobiSys | 8 |
| 2026 | AgentProg: Empowering Long-Horizon GUI Agents with Program-guided Context ManagementabstractThe rapid development of mobile GUI agents has stimulated growing research interest in long-horizon task automation. However, building agents for these tasks faces a critical bottleneck: the reliance on ever-expanding interaction history incurs substantial context overhead. Existing context management and compression techniques often fail to preserve vital semantic information, leading to degraded task performance. We propose AgentProg, a program-guided approach for agent context management that reframes the interaction history as a program with variables and control flow. By organizing information according to the structure of program, this structure provides a principled mechanism to determine which information should be retained and which can be discarded. We further integrate a global belief state mechanism inspired by Belief MDP framework to handle partial observability and adapt to unexpected environmental changes. Experiments on AndroidWorld and our extended long-horizon task suite demonstrate that AgentProg has achieved state-of-the-art success rates on these benchmarks. More importantly, it maintains robust performance on long-horizon tasks while baseline methods experience catastrophic degradation. Our system is open-sourced at https://github.com/MobileLLM/AgentProg. Shizuo Tian, Hao Wen 0004, Shanhui Zhao, Guohong Liu 0002, Ju Ren 0001, Yunxin Liu 0001, Yuanchun Li 0003 |
MobiSys | 8 |
| 2026 | Efficient Remote KV Cache Reuse with GPU-native Video Codec
Liang Mi, Weijun Wang 0001, Jinghan Chen, Ting Cao 0003, Haipeng Dai 0001, Yunxin Liu 0001 |
SIGCOMM | 6 |
| 2026 | LoAPE: A Load-Aware and Power-Elastic Platform for Green Serverless ComputingabstractEnergy efficiency in serverless computing has remained under-explored despite its growing adoption. To fill this gap, we propose LoAPE, a load-aware and power-elastic serverless platform. Unlike existing methods limited to CPU core frequency scaling or density-based consolidation, LoAPE leverages c-states and uncore frequency scaling, which offer greater energy benefits but pose deployment challenges. Uncore changes affect the performance of all co-located functions, and cstates require sustained idle periods for effective energy savings. LoAPE’s insight is to strategically exploit function elasticity. By creating idle intervals across cores and servers, it enables aggressive c-state activation and uncore frequency reduction. To achieve this goal, LoAPE adopts topology-aware instance scheduling and eviction, which scale the number of active servers along with serverless functions. Within servers, LoAPE proactively manages active cores based on incoming loads. Our evaluations demonstrate that LoAPE delivers a 3.02× improvement in cluster energy reduction compared to frequency-scaling frameworks and extends savings by 1.17× over state-of-the-art density-aware schedulers, while meeting performance requirements. Hanfei Geng, Yuanzhe Li 0001, Jichao Leng, Feng Zhao 0001, Yunxin Liu 0001 |
IEEE Trans. Computers | 5 |
| 2026 | MobiFuse: A High-Precision On-Device Depth Perception System With Multi-Data FusionabstractWe present MobiFuse, a high-precision depth perception system on mobile devices that combines dual RGB and Time-of-Flight (ToF) cameras. To achieve this, we leverage physical principles from various environmental factors to propose the Depth Error Indication (DEI) modality, characterizing the depth error of ToF and stereo-matching. Furthermore, we employ a progressive fusion strategy, merging geometric features from ToF and stereo depth maps with depth error features from the DEI modality to create precise depth maps. Additionally, we create a new ToF-Stereo depth dataset,RealToF, to train and validate our model. Our experiments demonstrate that MobiFuse excels over baselines by significantly reducing depth measurement errors by up to 77.7%. It also showcases strong generalization across diverse datasets and proves effectiveness in two downstream tasks: 3D reconstruction and 3D segmentation. The demo video of MobiFuse in real-life scenarios is available at the de-identified YouTube link. Tingting Long, Ju Ren 0001, Yunxin Liu 0001, Yudong Zhao, Yaoxue Zhang, Youngki Lee 0001 |
IEEE Trans. Mob. Comput. | 6 |
| 2026 | Bi-Level Bandwidth Coordination for Multiple Video Inference at the EdgeabstractHigh-definition (HD) cameras for surveillance and road traffic have experienced tremendous growth, demanding intensive computation resources for real-time analytics. Recently, offloading frames from the front-end device to the back-end edge server has shown great promise. In multi-stream competitive environments, efficient bandwidth management and proper scheduling are crucial to ensure both high inference accuracy and high throughput. To achieve this goal, we propose BiSwift, a bi-level framework that scales the concurrent real-time video analytics by a novel adaptive hybrid codec integrated with multi-level pipelines, and a global bandwidth controller for multiple video streams. The lower-level front-back-end collaborative mechanism (called adaptive hybrid codec) locally optimizes the accuracy and accelerates end-to-end video analytics for a single stream. The upper-level scheduler aims to accuracy fairness among multiple streams via the global bandwidth controller. The evaluation of BiSwift shows that BiSwift is able to real-time object detection on 9 streams with an edge device only equipped with an NVIDIA RTX3070 (8G) GPU. BiSwift improves 10%~21% accuracy and presents$1.2\sim 9\times $throughput compared with the state-of-the-art video analytics pipelines. Haipeng Dai 0001, Jinghan Chen, Liang Mi, Weijun Wang 0001, Yuanchun Li 0003, Tingting Yuan 0001, Yuben Qu, Yunxin Liu 0001, Xiaoming Fu 0001, Guihai Chen |
IEEE Trans. Netw. | 10 |
| 2025 | Empower Vision Applications with LoRA LMMabstractLarge Multimodal Models (LMMs) have shown significant progress in various complex vision tasks with the solid linguistic and reasoning capacity inherited from large language models (LMMs). Low-rank adaptation (LoRA) offers a promising method to integrate external knowledge into LMMs, compensating for their limitations on domain-specific tasks. However, the existing LoRA model serving is excessively computationally expensive and causes extremely high latency. In this paper, we present an end-to-end solution that empowers diverse vision tasks and enriches vision applications with LoRA LMMs. Our system, VaLoRA, enables accurate and efficient vision tasks by 1) an accuracy-aware LoRA adapter generation approach that generates LoRA adapters rich in domain-specific knowledge to meet application-specific accuracy requirements, 2) an adaptive-tiling LoRA adapters batching operator that efficiently computes concurrent heterogeneous LoRA adapters, and 3) a flexible LoRA adapter orchestration mechanism that manages application requests and LoRA adapters to achieve the lowest average response latency. We prototype VaLoRA on five popular vision tasks on three LMMs. Experiment results reveal that VaLoRA improves 24-62% of the accuracy compared to the original LMMs and reduces 20-89% of the latency compared to the state-of-the-art LoRA model serving systems. Liang Mi, Weijun Wang 0001, Wenming Tu, Qingfeng He, Xinyu Fang, Yazhu Dong, Yuanchun Li 0003, Meng Li 0010, Haipeng Dai 0001, Guihai Chen, Yunxin Liu 0001 |
EuroSys | 13 |
| 2025 | Data Center Cooling System Optimization Using Offline Reinforcement LearningabstractThe recent advances in information technology and artificial intelligence have fueled a rapid expansion of the data center (DC) industry worldwide, accompanied by an immense appetite for electricity to power the DCs. In a typical DC, around 30-40% of the energy is spent on the cooling system rather than on computer servers, posing a pressing need for developing new energy-saving optimization technologies for DC cooling systems. However, optimizing such real-world industrial systems faces numerous challenges, including but not limited to a lack of reliable simulation environments, limited historical data, and stringent safety and control robustness requirements. In this work, we present a novel physics-informed offline reinforcement learning (RL) framework for energy efficiency optimization of DC cooling systems. The proposed framework models the complex dynamical patterns and physical dependencies inside a server room using a purposely designed graph neural network architecture that is compliant with the fundamental time-reversal symmetry. Because of its well-behaved and generalizable state-action representations, the model enables sample-efficient and robust latent space offline policy learning using limited real-world operational data. Our framework has been successfully deployed and verified in a large-scale production DC for closed-loop control of its air-cooling units (ACUs). We conducted a total of 2000 hours of short and long-term experiments in the production DC environment. The results show that our method achieves 14-21% energy savings in the DC cooling system, without any violation of the safety or operational constraints. We have also conducted a comprehensive evaluation of our approach in a real-world DC testbed environment. Our results have demonstrated the significant potential of offline RL in solving a broad range of data-limited, safety-critical real-world industrial control problems. Xianyuan Zhan, Peng Cheng 0013, Ziteng He, Hanfei Geng, Jichao Leng, Huiwen Zheng, Tianshun Hong, Yunxin Liu 0001, Feng Zhao 0001 |
ICLR | 12 |
| 2025 | Demo: EdgeMind-OS: A Plug-and-Play Embodied Intelligence System for Real-Time On-Device DeploymentabstractBuilding an always-on, contextual AI assistant that proactively supports humans remains a central goal in Embodied AI—yet cloud-based pipelines struggle to meet due to delay, bandwidth, and privacy constraints. This demo presents EdgeMind-OS, a fully on-device intelligence system designed for embodied agents operating in real-world scenarios. Edge-Mind-OS features a hierarchical architecture combining a real-time StreamBrain, modular skill experts, and a dynamic scene-episode memory. Achieving up to 7.3× faster local processing, it enables low-latency, privacy-preserving, and plug-and-play deployment across tasks such as semantic navigation, spatial memory recall, and multimodal interaction. We demonstrate how EdgeMind-OS empowers a mobile robot with only basic locomotion capabilities to perform realtime, free-form user-robot interaction through autonomous perception, reasoning and action —without reliance on external cloud infrastructure. Jianyu Wei, Fucheng Jia, Liang Mi, Ruofei Ju, Xianye Wang, Yikai Zheng, Weijun Wang 0001, Shiqi Jiang 0002, Yunxin Liu 0001, Ting Cao 0003 |
MobiCom | 11 |
| 2025 | LLM-Explorer: Towards Efficient and Affordable LLM-based Exploration for Mobile AppsabstractLarge language models (LLMs) have opened new opportunities for automated mobile app exploration, an important and challenging problem that used to suffer from the difficulty of generating meaningful UI interactions. However, existing LLM-based exploration approaches rely heavily on LLMs to generate actions in almost every step, leading to a huge cost of token fees and computational resources. We argue that such extensive usage of LLMs is neither necessary nor effective, since many actions during exploration do not require, or may even be biased by the abilities of LLMs. Further, based on the insight that a precise and compact knowledge plays the central role for effective exploration, we introduce LLM-Explorer, a new exploration agent designed for efficiency and affordability. LLM-Explorer uses LLMs primarily for maintaining the knowledge instead of generating actions, and knowledge is used to guide action generation in a LLM-less manner. Based on a comparison with 5 strong baselines on 20 typical apps, LLM-Explorer was able to achieve the fastest and highest coverage among all automated app explorers, with over 148x lower cost than the state-of-the-art LLM-based approach. Shanhui Zhao, Hao Wen 0004, Wenjie Du 0004, Cheng Liang 0006, Yunxin Liu 0001, Xiaozhou Ye, Ye Ouyang, Yuanchun Li 0003 |
MobiCom | 5 |
| 2025 | AutoDroid-V2: Boosting SLM-based GUI Agents via Code GenerationabstractLarge language models (LLMs) have brought exciting new advances to mobile UI agents, a long-standing research field that aims to complete arbitrary natural language tasks through mobile UI interactions. However, existing UI agents usually demand powerful large language models that are difficult to be deployed locally on end-users' devices, raising huge concerns about user privacy and centralized serving cost. Inspired by the remarkable coding abilities of recent small language models (SLMs), we propose to convert the UI task automation problem to a code generation problem, which can be effectively solved by an on-device SLM and efficiently executed with an on-device code interpreter. Unlike normal coding tasks that can be extensively pre-trained with public datasets, generating UI automation code is challenging due to the diversity, complexity, and variability of target apps. Therefore, we adopt a document-centered approach that automatically builds fine-grained API documentation for each app and generates diverse task samples based on this documentation. By guiding the agent with the synthetic documents and task samples, it learns to generate precise and efficient scripts to complete unseen tasks. Based on detailed comparisons with state-of-the-art mobile UI agents, our approach effectively improves the mobile task automation with significantly higher success rates and lower latency/token consumption. Code is open-sourced at https://github.com/MobileLLM/AutoDroid-V2. Hao Wen 0004, Shizuo Tian, Borislav Pavlov, Wenjie Du 0004, Ge Chang 0002, Shanhui Zhao, Yunxin Liu 0001, Ya-Qin Zhang, Yuanchun Li 0003 |
MobiSys | 9 |
| 2025 | Region-based Content Enhancement for Efficient Video Analytics at the Edge
Weijun Wang 0001, Liang Mi, Shaowei Cen, Haipeng Dai 0001, Yuanchun Li 0003, Xiaoming Fu 0001, Yunxin Liu 0001 |
NSDI | 7 |
| 2025 | A goal-oriented document-grounded dialogue based on evidence generation
Yong Song 0003, Hongjie Fan, Yunxin Liu 0001, Xiaozhou Ye, Ye Ouyang |
Data Knowl. Eng. | 4 |
| 2025 | Serving MoE Models on Resource-Constrained Edge Devices via Dynamic Expert SwappingabstractMixture of experts (MoE) is a popular technique in deep learning that improves model capacity with conditionally-activated parallel neural network modules (experts). However, serving MoE models in resource-constrained latency-critical edge scenarios is challenging due to the significantly increased model size and complexity. In this paper, we first analyze the behavior pattern of MoE models in continuous inference scenarios, which leads to three key observations about the expert activations, including temporal locality, exchangeability, and skippable computation. Based on these observations, we introduce PC-MoE, an inference framework for resource-constrained continuous MoE model serving. The core of PC-MoE is a new data structure,Parameter Committee, that intelligently maintains a subset of important experts in use to reduce resource consumption. To evaluate the effectiveness of PC-MoE, we conduct experiments using state-of-the-art MoE models on common computer vision and natural language processing tasks. The results demonstrate optimal trade-offs between resource consumption and model accuracy achieved by PC-MoE. For instance, on object detection tasks with the Swin-MoE model, our approach can reduce memory usage and latency by 42.34% and 18.63% with only 0.10% accuracy degradation. Yuanchun Li 0003, Weijun Wang 0001, Linghe Kong, Yunxin Liu 0001 |
IEEE Trans. Computers | 5 |
| 2025 | DSTC: Dual-Side Sparse Tensor Core for DNNs Acceleration on Modern GPU ArchitecturesabstractLeveraging sparsity in deep neural network (DNN) models holds significant promise for accelerating model inference. However, current GPUs can only harness sparsity in model weights, leaving activations unutilized due to their dynamic and unpredictable nature, which poses a considerable challenge for exploitation. In our research, we introduce a novel architectural approach aimed at effectively leveraging dual-side sparsity, encompassing both weight and activation sparsity. Our methodology involves a systematic examination of previous sparsity-related architectures, and culminating in the proposal of an uncharted paradigm that combines outer-product computation primitive and bitmap-based encoding format. Our approach showcases feasibility through minimal modifications to existing production-scale inner-product-based Tensor Cores. We introduce a set of innovative ISA extensions and carefully co-design matrix-matrix multiplication and convolution algorithms, the two predominant computation patterns in contemporary DNN models, to exploit our novel dual-side sparse Tensor Core. Our evaluation demonstrates the efficacy of our design, unlocking the full potential of dual-side DNN sparsity and delivering performance enhancements of up to an order of magnitude while incurring only modest hardware overhead. Chen Zhang 0001, Yang Wang 0053, Cong Guo 0003, Yunxin Liu 0001, Jingwen Leng, Zhigang Ji, Yuan Xie 0001, Ru Huang 0001 |
IEEE Trans. Computers | 5 |
| 2025 | Fine-Grained Structured Sparse Computing for FPGA-Based AI InferenceabstractWith the explosive growth in the number of parameters in deep neural networks (DNNs), sparsity-centric algorithm and hardware designs have become critical for low-latency AI serving systems. However, the inherent randomness in pruning methods often leads to fragmented data access and irregular computation patterns in sparse matrices, resulting in significantly reduced hardware efficiency. Addressing the balance between the ‘randomness’ required to maintain model accuracy and the ‘regularity’ needed for efficient hardware design is crucial for realizing effective sparse computing in AI. This article proposes a fine-grained structured sparsity (FSS) paradigm. The pruned sparse matrices in this paradigm exhibit characteristics of ‘local randomness’ and ‘global regularity’. This dual-feature design allows AI accelerator hardware based on the FSS paradigm to maintain both high model accuracy and efficient hardware design. We implemented this novel accelerator on the Xilinx Alveo U280 and validated our concept across three different AI models, including CNN, RNN, and LLM, demonstrating performance that significantly outperforms prior methods. Chen Zhang 0001, Shijie Cao, Guohao Dai 0001, Chenbo Geng, Zhuliang Yao, Wencong Xiao, Yunxin Liu 0001, Ming Wu 0007, Guangyu Sun 0003, Zhigang Ji, Runsheng Wang, Ru Huang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2025 | TimelyNet: Adaptive Neural Architecture for Autonomous Driving with Dynamic DeadlineabstractTo maintain driving safety, the execution of neural network-based autonomous driving pipelines must meet the dynamic deadlines in response to the changing environment and vehicle’s velocity. To this end, this article proposes a real-time neural architecture adaptation approach, called TimelyNet, which uses a supernet to replace the most compute-intensive neural network module in an existing end-to-end autonomous driving pipeline. From the supernet, TimelyNet samples subnets with varying inference latency levels to meet the dynamic deadlines during run-time driving without fine-tuning. Specifically, TimelyNet employs a one-shot prediction method that jointly uses a lookup table and an invertible neural network to periodically determine the optimal hyperparameters of a subnet to meet its execution deadline while achieving the highest possible accuracy. The lookup table stores multiple subnet architectures with different latencies, while the invertible neural network models the distribution of the optimal subnet architecture given the latency. Extensive evaluation based on hardware-in-the-loop CARLA simulations shows that TimelyNet-integrated driving pipelines achieve the best driving safety, characterized by the lowest wrong-lane driving rate and zero collisions, compared with several baselines, including the state-of-the-art driving pipelines. Duc Van Le, Yuanchun Li 0003, Yunxin Liu 0001, Rui Tan 0001 |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2025 | Squeezer: Efficient Multi-DNN Inference for Edge Video Analytics via Cross-Model SchedulingabstractVideo analytics at the edge is becoming increasingly prevalent in many scenarios, such as smart campuses and intelligent factories. These applications often consist of multiple subtasks, which necessitates the optimization for multi-DNN (Deep Neural Network) inference. Due to limited consideration over cross-model scheduling, current practices cannot fully leverage available computing resources, leading to suboptimal performance. To address this, we propose Squeezer, a multiDNN serving framework that holistically schedules multiple DNN models on an edge server with a single GPU. Squeezer decouples the cross-model scheduling into a two-layered approach, which involves (1) balanced operator grouping which partitions operators of multiple DNN models into groups, significantly reducing the scheduling complexity and (2) kernel scheduler which orchestrates parallel execution within each group by considering the interplay among kernels running in parallel, thereby enabling cross-model optimizations in multi-DNN inference. Performance evaluation results demonstrate that Squeezer outperforms state-of-the-art baselines, achieving up to 1.91× improvement in system throughput. Lingxiao Ma, Ziyan Fu 0001, Yuanchun Li 0003, Ju Ren 0001, Yaoxue Zhang, Yunxin Liu 0001 |
IEEE Trans. Mob. Comput. | 8 |
| 2025 | AdaEvo: Edge-Assisted Continuous and Timely DNN Model Evolution for Mobile DevicesabstractMobile video applications today have attracted significant attention. Deep learning model (e.g., deep neural network, DNN) compression is widely used to enable on-device inference for facilitating robust and private mobile video applications. The compressed DNN, however, is vulnerable to the agnostic data drift of the live video captured from the dynamically changing mobile scenarios. To combat the data drift, mobile ends rely on edge servers to continuously evolve and re-compress the DNN with freshly collected data. We design a framework, AdaEvo, that efficiently supports the resource-limited edge server handling mobile DNN evolution tasks from multiple mobile ends. The key goal of AdaEvo is to maximize the average quality of experience (QoE), i.e., the proportion of high-quality DNN service time to the entire life cycle, for all mobile ends. Specifically, it estimates the DNN accuracy drops at the mobile end without labels and performs a dedicated video frame sampling strategy to control the size of retraining data. In addition, it balances the limited computing and memory resources on the edge server and the competition between asynchronous tasks initiated by different mobile users. With an extensive evaluation of real-world videos from mobile scenarios and across four diverse mobile tasks, experimental results show that AdaEvo enables up to 34% accuracy improvement and 32% average QoE improvement. Lehao Wang, Zhiwen Yu 0001, Haoyi Yu, Sicong Liu 0005, Yaxiong Xie, Bin Guo 0001, Yunxin Liu 0001 |
IEEE Trans. Mob. Comput. | 7 |
| 2025 | AdaWiFi, Collaborative WiFi Sensing for Cross-Environment AdaptationabstractDeep learning (DL) based Wi-Fi sensing has witnessed great development in recent years. Although decent results have been achieved in certain scenarios, Wi-Fi based activity recognition is still difficult to deploy in real smart homes due to the limited cross-environment adaptability, i.e. a well-trained Wi-Fi sensing neural network in one environment is hard to adapt to other environments. To address this challenge, we proposeAdaWiFi, a DL-based Wi-Fi sensing framework that allows multiple Internet-of-Things (IoT) devices to collaborate and adapt to various environments effectively. The key innovation ofAdaWiFiincludes a collective sensing model architecture that utilizes complementary information between distinct devices and avoids the biased perception of individual sensors and an accompanying model adaptation technique that can transfer the sensing model to new environments with limited data. We evaluate our system on a public dataset and a custom dataset collected from three complex sensing environments. The results demonstrate thatAdaWiFiis able to achieve significantly better sensing adaptation effectiveness (e.g. 30% higher accuracy with one-shot adaptation) as compared with state-of-the-art baselines. Naiyu Zheng, Yuanchun Li 0003, Shiqi Jiang 0002, Yuanzhe Li 0001, Rongchun Yao, Chuchu Dong, Zhimeng Yin 0001, Yunxin Liu 0001 |
IEEE Trans. Mob. Comput. | 10 |
| 2024 | SwapMoE: Serving Off-the-shelf MoE-based Large Language Models with Tunable Memory BudgetabstractRui Kong, Yuanchun Li, Qingtian Feng, Weijun Wang, Xiaozhou Ye, Ye Ouyang, Linghe Kong, Yunxin Liu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Yuanchun Li 0003, Qingtian Feng, Weijun Wang 0001, Xiaozhou Ye, Ye Ouyang, Linghe Kong, Yunxin Liu 0001 |
ACL (1) | 8 |
| 2024 | Amanda: Unified Instrumentation Framework for Deep Neural NetworksabstractThe success of deep neural networks (DNNs) has sparked efforts to analyze (e.g., tracing) and optimize (e.g., pruning) them. These tasks have specific requirements and ad-hoc implementations in current execution backends like TensorFlow/PyTorch, which require developers to manage fragmented interfaces and adapt their codes to diverse models. In this study, we propose a new framework called Amanda to streamline the development of these tasks. We formalize the implementation of these tasks as neural network instrumentation, which involves introducing instrumentation into the operator level of DNNs. This allows us to abstract DNN analysis and optimization tasks as instrumentation tools on various DNN models. We build Amanda with two levels of APIs to achieve a unified, extensible, and efficient instrumentation design. The user-level API provides a unified operator-grained instrumentation API for different backends. Meanwhile, internally, we design a set of callback-centric APIs for managing and optimizing the execution of original and instrumentation codes in different backends. Through these design principles, the Amanda framework can accommodate a broad spectrum of use cases, such as tracing, profiling, pruning, and quantization, across different backends (e.g., TensorFlow/PyTorch) and execution modes (graph/eager mode). Moreover, our efficient execution management ensures that the performance overhead is typically kept within 5%. Yue Guan 0003, Yuxian Qiu, Jingwen Leng, Fan Yang 0024, Shuo Yu 0006, Yunxin Liu 0001, Yu Feng 0007, Yuhao Zhu 0001, Lidong Zhou, Yun Liang 0001, Chen Zhang 0001, Chao Li 0009, Minyi Guo |
ASPLOS (1) | 6 |
| 2024 | TESLA: Thermally Safe, Load-Aware, and Energy-Efficient Cooling Control System for Data CentersabstractThe increasing demand for artificial intelligence and cloud computing has led to skyrocketing energy consumption of data centers (DCs). This paper focuses on tackling this energy challenge through cooling control system optimization, which aims to ensure thermal safety with minimal cooling energy consumption. Current industry practice involves human operators, while many data-driven methods have also been proposed. However, human intervention often results in unnecessary energy consumption, particularly in the face of fluctuating server loads, whereas existing data-driven methods struggle to maintain thermal safety in practice. To overcome these issues, we propose TESLA, a thermally safe, load-aware, and energy-efficient cooling control system for data centers. TESLA employs a novel data-driven framework that integrates domain knowledge to predict DC temperature and cooling energy under dynamic server load. Based on these predictions, a Bayesian optimizer (BO) finds the energy-optimal settings for the cooling system at every control step. Besides cooling energy, BO’s optimization objective also includes minimizing cooling interruption that causes rapid temperature rise within the data center and leads to thermal safety violations. We deploy TESLA on a real data-center testbed and show that it achieves on average <?TeX $10.1\%$?> Math 1 cooling energy saving relative to a fixed cooling system parameter setting and no thermal safety violation relative to previous data-driven methods. Hanfei Geng, Yuanzhe Li 0001, Jichao Leng, Xianyuan Zhan, Yuanchun Li 0003, Feng Zhao 0001, Yunxin Liu 0001 |
ICPP | 9 |
| 2024 | BiSwift: Bandwidth Orchestrator for Multi-Stream Video Analytics on EdgeabstractHigh-definition (HD) cameras for surveillance and road traffic have experienced tremendous growth, demanding intensive computation resources for real-time analytics. Recently, offloading frames from the front-end device to the back-end edge server has shown great promise. In multi-stream competitive environments, efficient bandwidth management and proper scheduling are crucial to ensure both high inference accuracy and high throughput. To achieve this goal, we propose BiSwift, a bi-level framework that scales the concurrent real-time video analytics by a novel adaptive hybrid codec integrated with multi-level pipelines, and a global bandwidth controller for multiple video streams. The lower-level front-back-end collaborative mechanism (called adaptive hybrid codec) locally optimizes the accuracy and accelerates end-to-end video analytics for a single stream. The upper-level scheduler aims to accuracy fairness among multiple streams via the global bandwidth controller. The evaluation of BiSwift shows that BiSwift is able to real-time object detection on 9 streams with an edge device only equipped with an NVIDIA RTX3070 (8G) GPU. BiSwift improves 10%∼21% accuracy and presents 1.2∼ 9× throughput compared with the state-of-the-art video analytics pipelines. Weijun Wang 0001, Tingting Yuan 0001, Liang Mi, Haipeng Dai 0001, Yunxin Liu 0001, Xiaoming Fu 0001 |
INFOCOM | 6 |
| 2024 | AutoDroid: LLM-powered Task Automation in AndroidabstractMobile task automation is an attractive technique that aims to enable voice-based hands-free user interaction with smartphones. However, existing approaches suffer from poor scalability due to the limited language understanding ability and the non-trivial manual efforts required from developers or endusers. The recent advance of large language models (LLMs) in language understanding and reasoning inspires us to rethink the problem from a model-centric perspective, where task preparation, comprehension, and execution are handled by a unified language model. In this work, we introduce AutoDroid, a mobile task automation system capable of handling arbitrary tasks on any Android application without manual efforts. The key insight is to combine the commonsense knowledge of LLMs and domain-specific knowledge of apps through automated dynamic analysis. The main components include a functionality-aware UI representation method that bridges the UI with the LLM, exploration-based memory injection techniques that augment the app-specific domain knowledge of LLM, and a multi-granularity query optimization module that reduces the cost of model inference. We integrate AutoDroid with off-the-shelf LLMs including online GPT-4/GPT-3.5 and on-device Vicuna, and evaluate its performance on a new benchmark for memory-augmented Android task automation with 158 common tasks. The results demonstrated that AutoDroid is able to precisely generate actions with an accuracy of 90.9%, and complete tasks with a success rate of 71.3%, outperforming the GPT-4-powered baselines by 36.4% and 39.7%. Hao Wen 0004, Yuanchun Li 0003, Guohong Liu 0002, Shanhui Zhao, Toby Jia-Jun Li, Shiqi Jiang 0002, Yunhao Liu 0001, Yunxin Liu 0001 |
MobiCom | 10 |
| 2024 | FlexNN: Efficient and Adaptive DNN Inference on Memory-Constrained Edge DevicesabstractDue to the popularity of deep neural networks (DNNs) and considerations over network overhead, data privacy, and inference latency, there is a growing interest in deploying DNNs to edge devices in recent years. However, the limited memory becomes a major bottleneck for on-device DNN deployment, making it crucial to reduce the memory footprint of DNN. The mainstream model customization solutions require intensive deployment efforts and may lead to severe accuracy degradation, and existing deep learning (DL) frameworks don't take memory as a priority. Besides, recent works to enhance the memory management scheme cannot be directly applied because of several challenges, including the unbalanced memory footprint across layers, the inevitable overhead of memory management, and the memory budget dynamicity. To tackle these challenges, we introduce FlexNN, an efficient and adaptive memory management framework for DNN inference on memory-constrained devices. FlexNN uses a slicing-loading-computing joint planning approach, to achieve optimal memory utilization and minimal memory management overhead. We implemented FlexNN atop NCNN, and conducted comprehensive evaluations with common model architectures on various devices. The results have shown that our approach is able to adapt to different memory constraints with optimal latency-memory trade-offs. For example, FlexNN can reduce the memory consumption by 93.81% with only a 3.64% increase in latency, as compared with the original NCNN on smartphones. Yuanchun Li 0003, Yuanzhe Li 0001, Ting Cao 0003, Yunxin Liu 0001 |
MobiCom | 5 |
| 2024 | Poster: Enabling Agent-centric Interaction on Smartphones with LLM-based UI ReassemblingabstractIn this poster, we introduce a novel dynamic user interface (UI) specifically designed for mobile devices powered by large language models (LLMs) agents. The advent of LLMs has led to a surge in deploying LLM-based agents on personal and Internet of Things (IoT) devices, with the aim of facilitating various daily tasks through device manipulation. However, this integration poses a significant challenge: how to intelligently and flexibly select and present information both during and after the execution of tasks, ensuring users are well-informed about the operations and can access the desired results conveniently. To address this challenge, we propose a UI reassembling method. This method allows for analyzing and strategically combining different mobile applications and their UI components, enabling the dynamic construction and adjustment of UIs tailored to user needs. Our prototype exhibits promising performance, with the UI selection module achieving an F1 score of 0.74. This innovative approach opens up exciting possibilities of new user-device interaction paradigm, leveraging the capabilities of LLMs to enhance the user experience in handling mobile and IoT devices. Hao Wen 0004, Wenjie Du 0004, Yuanchun Li 0003, Yunxin Liu 0001 |
MobiSys | 4 |
| 2024 | Empowering In-Browser Deep Learning Inference on Edge Through Just-In-Time Kernel OptimizationabstractWeb is increasingly becoming the primary platform to deliver AI services onto edge devices, making in-browser deep learning (DL) inference more prominent. Nevertheless, the heterogeneity of edge devices, combined with the underdeveloped state of Web hardware acceleration practices, hinders current in-browser inference from achieving its full performance potential on target devices. Fucheng Jia, Shiqi Jiang 0002, Ting Cao 0003, Tianrui Xia, Yuanchun Li 0003, Qipeng Wang 0001, Ju Ren 0001, Yunxin Liu 0001, Lili Qiu, Mao Yang 0004 |
MobiSys | 11 |
| 2024 | TIM: Enabling Large-Scale White-Box Testing on In-App Deep Learning ModelsabstractIntelligent Applications (iApps), equipped with in-App deep learning (DL) models, are emerging to provide reliable DL inference services. However, in-App DL models are typically compiled into inference-only versions to enhance system performance, thereby impeding the evaluation of DL models. Specifically, the assessment of in-App models currently relies on black-box testing methods rather than direct white-box testing approaches. In this work, we propose TIM, an automated tool designed for conducting large-scale white-box testing of in-App models. Taking an iApp as input, TIM can lift the black-box (i.e., inference-only) in-App DL model into a backpropagation-enabled one and package it together, allowing comprehensive DL model testing or security issues detection. TIM proposes two reconstruction techniques to convert the inference-only model to a backpropagation-enabled version and reconstruct the DL-related IO processing code. In our experiments, we utilize TIM to extract 100 unique commercial in-App models and convert the models to white-box models, enabling backpropagation functionality. Experimental results show that TIM’s reconstruction techniques exhibit high accuracy. We open-source our prototype and part of the experimental data on the websitehttps://zenodo.org/record/7548141. Hao Wu 0067, Yuhang Gong, Xiaopeng Ke, Hanzhong Liang, Fengyuan Xu, Yunxin Liu 0001, Sheng Zhong 0002 |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2024 | Seamless Cross-Edge Service Migration for Real-Time Rendering ApplicationsabstractSeamless cross-edge migration for real-time rendering applications is challenging. The strong interactive nature of real-time rendering applications demands a downtime lower than 15ms to achieve an imperceptible migration. Existing methods based on virtual machine migration and container migration suffer from unpleasant downtime brought by dirty page retransmission-induced repeated memory data copy and the shared storage failure-induced extensive disk data copy. In this paper, we propose Cloud-assisted Service Migration (CSM) which leverages cloud-edge collaboration to achieve seamless service migration for real-time rendering applications. CSM improves service migration user experience in three folds: First, it introduces a dual rendering mechanism to bypass the peer-to-peer data copy and compresses the freezing stage. Second, a user equipment-centric session switch mechanism is proposed to save time by well coordinating application session switches and 5G user plane session switches. Third, a smooth switching mechanism is leveraged to prevent unpleasant frame flickers during session switching. We implement CSM in edge-rendering multiplayer games and deploy it on a 5G test bed with a full-stack user plane protocol stack. The evaluation results show that CSM can reduce downtime to < 14ms and the service migration process is user imperceptible. Yuanzhe Li 0001, Shangguang Wang, Yuanchun Li 0003, Ao Zhou 0001, Mengwei Xu 0001, Xiao Ma 0009, Yunxin Liu 0001 |
IEEE Trans. Mob. Comput. | 7 |
| 2024 | FLASH: Heterogeneity-Aware Federated Learning at ScaleabstractFederated learning (FL) becomes a promising machine learning paradigm. The impact of heterogeneous hardware specifications and dynamic states on the FL process has not yet been studied systematically. This paper presents the first large-scale study of this impact based on real-world data collected from 136k smartphones. We conducted extensive experiments on our proposed heterogeneity-aware FL platform namelyFLASH, to systematically explore the performance of state-of-the-art FL algorithms and key FL configurations in heterogeneity-aware and -unaware settings, finding the following. (1) Heterogeneity causes accuracy to drop by up to 9.2% and convergence time to increase by 2.32×. (2) Heterogeneity negatively impacts popular aggregation algorithms, e.g., the accuracy variance reduction brought byq-FedAvgdrops by 17.5%. (3) Heterogeneity does not worsen the accuracy loss caused by gradient-compression algorithms significantly, but it compromises the convergence time by up to 2.5×. (4) Heterogeneity hinders client-selection algorithms from selecting wanted clients, thus reducing effectiveness. e.g., the accuracy increase brought by the state-of-the-art client-selection algorithm drops by 73.9%. (5) Heterogeneity causes the optimal FL hyper-parameters to drift significantly. More specifically, the heterogeneity-unaware setting favors looser deadline and higher reporting fraction to achieve better training performance. (6) Heterogeneity results in non-trivial failed clients (more than 10%) and leads to participation bias (the top 30% of clients contribute 86% of computations). Our FLASH platform and data have been publicly open sourced. Chengxu Yang, Mengwei Xu 0001, Qipeng Wang 0001, Zhenpeng Chen 0001, Yun Ma 0002, Kaigui Bian, Gang Huang 0001, Yunxin Liu 0001, Xin Jin 0008, Xuanzhe Liu |
IEEE Trans. Mob. Comput. | 9 |
| 2024 | HiMoDepth: Efficient Training-Free High-Resolution On-Device Depth PerceptionabstractHigh-resolution depth estimation, with a minimum resolution of$1280\times 960$, is essential for achieving more immersive experiences in on-device 3D vision applications. However, implementing high-resolution solutions on resource-limited mobile devices presents significant challenges, such as the need for additional expensive depth sensors, computation-intensive machine learning models requiring large-scale datasets, or the need for device motion while the target object remains stationary. In this study, we propose HiMoDepth, an efficient training-free high-resolution depth estimation system that utilizes widely-available on-device dual cameras. HiMoDepth consists of two modules: 1) homogenizing the on-device heterogeneous cameras by iteratively cropping the Field-of-Views to make the focal length of the cameras equal and filtering out the out-of-sync frames based on time stamps, and 2) designing a hierarchical mobile GPU-friendly stereo matching method that effectively reduces the latency of stereo matching with high-resolution depth maps by using efficient data layout, reducing the number of memory accesses, and searching the corresponding pixel over a coarse-to-fine hierarchy. We implement HiMoDepth on multiple commodity mobile devices and conduct comprehensive evaluations. Experimental results show that HiMoDepth significantly outperforms the baselines in both accuracy and running speed on mobile devices that support high-resolution depth maps. Ju Ren 0001, Bangwen He, Youngki Lee 0001, Ting Cao 0003, Yuanchun Li 0003, Yaoxue Zhang, Yunxin Liu 0001 |
IEEE Trans. Mob. Comput. | 10 |
| 2023 | FedTherapist: Mental Health Monitoring with User-Generated Linguistic Expressions on Smartphones via Federated LearningabstractPsychiatrists diagnose mental disorders via the linguistic use of patients.Still, due to data privacy, existing passive mental health monitoring systems use alternative features such as activity, app usage, and location via mobile devices.We propose FedTherapist, a mobile mental health monitoring system that utilizes continuous speech and keyboard input in a privacy-preserving way via federated learning.We explore multiple model designs by comparing their performance and overhead for FedTherapist to overcome the complex nature of on-device language model training on smartphones.We further propose a Context-Aware Language Learning (CALL) methodology to effectively utilize smartphones' large and noisy text for mental health signal sensing.Our IRBapproved evaluation of the prediction of selfreported depression, stress, anxiety, and mood from 46 participants shows higher accuracy of FedTherapist compared with the performance with non-language features, achieving 0.15 AU-ROC improvement and 8.21% MAE reduction. Jaemin Shin 0005, Hyungjun Yoon, Seungjoo Lee, Yunxin Liu 0001, Jinho D. Choi, Sung-Ju Lee 0001 |
EMNLP | 5 |
| 2023 | The First Decade of Computing and Network ConvergenceabstractThe past few years have witnessed an exciting journey from cloud-network convergence to Computing and Network Convergence (CNC) across major telcos worldwide. The pandemic in the past 3 years has further accelerated the digital transformation for verticals, which present high demands on both connectivity and computing. This paper reviews the development progression of CNC by Standards Developing Organizations (SDOs) and surveys major telco's strategies evolving from cloud-network convergence to CNC. The authors propose a new scheme of CNC to operate and manage CNC infrastructure and services. In particular, CNC Brain, as the core component of CNC, is proposed to collaborate, orchestrate, schedule, and allocate computing resources across cloud-edge-end through networks to serve any computing task anywhere and anytime. The authors also design a novel algorithm portfolio for CNC Brain to leverage least cost of CNC resources and optimal routes to meet the given service requirements of CNC. A CNC testbed is implemented with full CNC functionalities and experiments are illustrated for verification of the CNC scheme. The paper also forward looks at the roadmap with the critical milestones of CNC for the first decade. Ye Ouyang, Xiaozhou Ye, Yunxin Liu 0001 |
ICC | 4 |
| 2023 | OliVe: Accelerating Large Language Models via Hardware-friendly Outlier-Victim Pair QuantizationabstractTransformer-based large language models (LLMs) have achieved great success with the growing model size. LLMs' size grows by 240× every two years, which outpaces the hardware progress and makes model inference increasingly costly. Model quantization is a promising approach to mitigate the widening gap between LLM size and hardware capacity. However, the existence of outliers, values with significant magnitudes, in LLMs makes existing quantization methods less effective. Prior outlier-aware quantization schemes adopt sparsity encoding techniques to separate outliers from normal values where the process requires global coordination (e.g., a global sparsity coordination list). This incurs complex encoding/decoding hardware logics and an extra orchestration controller for the computation between outlier and normal values. As such, it is not hardware-efficient and hence only achieves sub-optimal quantization benefits. Cong Guo 0003, Weiming Hu 0005, Jingwen Leng, Chen Zhang 0001, Fan Yang 0024, Yunxin Liu 0001, Minyi Guo, Yuhao Zhu 0001 |
ISCA | 7 |
| 2023 | PatchBackdoor: Backdoor Attack against Deep Neural Networks without Model ModificationabstractBackdoor attack is a major threat to deep learning systems in safety-critical scenarios, which aims to trigger misbehavior of neural network models under attacker-controlled conditions. However, most backdoor attacks have to modify the neural network models through training with poisoned data and/or direct model editing, which leads to a common but false belief that backdoor attack can be easily avoided by properly protecting the model. In this paper, we show that backdoor attacks can be achieved without any model modification. Instead of injecting backdoor logic into the training data or the model, we propose to place a carefully-designed patch (namely backdoor patch) in front of the camera, which is fed into the model together with the input images. The patch can be trained to behave normally at most of the time, while producing wrong prediction when the input image contains an attacker-controlled trigger object. Our main techniques include an effective training method to generate the backdoor patch and a digital-physical transformation modeling method to enhance the feasibility of the patch in real deployments. Extensive experiments show that PatchBackdoor can be applied to common deep learning models (VGG, MobileNet, ResNet) with an attack success rate of 93% to 99% on classification tasks. Moreover, we implement PatchBackdoor in real-world scenarios and show that the attack is still threatening. Yizhen Yuan, Shenghao Xie 0002, Yuanchun Li 0003, Yunxin Liu 0001 |
ACM Multimedia | 5 |
| 2023 | A Proposal-Improved Few-Shot Embedding Model with Contrastive Learning
Fucai Gong, Yunxin Liu 0001, Xiaozhou Ye, Ye Ouyang |
MMM (2) | 5 |
| 2023 | LUT-NN: Empower Efficient Neural Network Inference with Centroid Learning and Table LookupabstractOn-device Deep Neural Network (DNN) inference consumes significant computing resources and development efforts. To alleviate that, we propose LUT-NN, the first system to empower inference by table lookup, to reduce inference cost. LUT-NN learns the typical features for each operator, named centroid, and precompute the results for these centroids to save in lookup tables. During inference, the results of the closest centroids with the inputs can be read directly from the table, as the approximated outputs without computations. Xiaohu Tang 0003, Yang Wang 0053, Ting Cao 0003, Li Lyna Zhang, Qi Chen 0009, Deng Cai 0001, Yunxin Liu 0001, Mao Yang 0004 |
MobiCom | 7 |
| 2023 | AdaptiveNet: Post-deployment Neural Architecture Adaptation for Diverse Edge EnvironmentsabstractDeep learning models are increasingly deployed to edge devices for real-time applications. To ensure stable service quality across diverse edge environments, it is highly desirable to generate tailored model architectures for different conditions. However, conventional pre-deployment model generation approaches are not satisfactory due to the difficulty of handling the diversity of edge environments and the demand for edge information. In this paper, we propose to adapt the model architecture after deployment in the target environment, where the model quality can be precisely measured and private edge data can be retained. To achieve efficient and effective edge model generation, we introduce a pretraining-assisted on-cloud model elastification method and an edge-friendly on-device architecture search method. Model elastification generates a high-quality search space of model architectures with the guidance of a developer-specified oracle model. Each subnet in the space is a valid model with different environment affinity, and each device efficiently finds and maintains the most suitable subnet based on a series of edge-tailored optimizations. Extensive experiments on various edge devices demonstrate that our approach is able to achieve significantly better accuracy-latency tradeoffs (e.g. 46.74% higher on average accuracy with a 60% latency budget) than strong baselines with minimal overhead (13 GPU hours in the cloud and 2 minutes on the edge server). Hao Wen 0004, Yuanchun Li 0003, Zunshuai Zhang, Shiqi Jiang 0002, Xiaozhou Ye, Ye Ouyang, Yunxin Liu 0001 |
MobiCom | 8 |
| 2023 | NN-Stretch: Automatic Neural Network Branching for Parallel Inference on Heterogeneous Multi-ProcessorsabstractMobile devices are increasingly equipped with heterogeneous multiprocessors, e.g., CPU + GPU + DSP. Yet existing Neural Network (NN) inference fails to fully utilize the computing power of the heterogeneous multi-processors due to the sequential structures of NN models. Towards this end, this paper proposes NN-Stretch, a new model adaption strategy, as well as the supporting system. It automatically branches a given model according to the processor architecture characteristics. Compared to other popular model adaption techniques such as model pruning that often sacrifices accuracy, NN-Stretch accelerates inference while preserving accuracy. Jianyu Wei, Ting Cao 0003, Shijie Cao, Shiqi Jiang 0002, Shaowei Fu, Mao Yang 0004, Yanyong Zhang, Yunxin Liu 0001 |
MobiSys | 8 |
| 2023 | Understanding the Impact of Quantum Noise on Quantum ProgramsabstractQuantum computing is expected to introduce the next era of computing speed and power, and its software - quantum program is gaining increasing research interest in the software engineering community. A significant characteristic of quantum computing is the existence of noise. Unlike classical computers where the output of a program is usually deterministic, the execution of a quantum program may be affected by quantum noise. Such a difference may cause difficulties or misunderstandings for developers shifting from classical programming to quantum programming. To understand the impact of quantum noise on quantum programs and its implications for software developers, we conduct a series of studies with real-world quantum programs and quantum computing environments. Specifically, we first measure and analyze the noise in a real quantum computer by testing it with a basic quantum program. We find that a non-neglectable amount of quantum noise generally exists in real quantum computers. Then we investigate the robustness of quantum programs against different quantum noises by testing 18 real-world quantum programs and 50,000 randomly generated quantum circuits in simulated and real environments. We observe that quantum noise can significantly influence the correctness of quantum programs, and different quantum circuit structures show diverse sensitivity patterns under the same noise. Based on the observations, we build a machine learning model to predict the fidelity of a quantum program under certain quantum noise. The model achieves a small average fidelity prediction error, meaning the impact of noise can be precisely estimated statistically. Zhonghao Pan, Yang Feng 0003, Yunxin Liu 0001, Yuanchun Li 0003 |
SANER | 4 |
| 2023 | Environment-aware Testing for DNN-based Smart-home WiFi Sensing SystemsabstractWiFi-based human activity recognition is a promising sensing application in smart homes due to the low cost, wide availability, and privacy preservation of WiFi devices. However, pushing WiFi sensing technology to industry-scale deployment is difficult due to its poor robustness against environment differences. How to systematically test such sensing system is crucial to improve its practicality, and is also challenging because the sensing performance is significantly influenced by the underlying physical environments. In this paper, we introduce the problem of testing environment-dependent sensing systems, including how to measure test coverage and how to effectively generate data to improve the coverage. We describe our initial attempts on examining test sufficiency with environment-neuron joint coverage and improving the coverage through targeted environment variations and signal transformations. Our experiments have demonstrated the higher effectiveness of using environment-neuron coverage to represent test sufficiency, as compared with using the conventional neuron coverage. Meanwhile, the coverage-guided sensing data generation can lead to higher accuracy of the sensing system under changing environments. Naiyu Zheng, Chuchu Dong, Yuanzhe Li 0001, Yunxin Liu 0001, Yuanchun Li 0003 |
SANER | 6 |
| 2023 | ${{\sf S \text{-}UbiTap}}$S-UbiTap: Leveraging Acoustic Dispersion for Ubiquitous and Scalable Touch Interface on Solid SurfacesabstractAs various computing devices, such as smartphones, IoT devices, smart speakers etc, becomes omnipresent in our daily lives, interest in ubiquitous computing interfaces is increasing. In response to this, various studies have introduced on-surface input techniques that leverage the surface of surrounding objects as touch interfaces. However, most of them struggle to support ubiquitous interaction due to their dependency on specific hardware or environments. In this work, we propose${{\sf S \text{-}UbiTap}}$, an input method that turns any flat solid surface into a touch input space by listening to sound (i.e., with microphones already present in the commodity devices). More specifically, we develop a novel touch localization technique that leverages the physical phenomenon, calleddispersion, which is the characteristic of sound as it travels through solid surfaces, and address the challenges that limit existing acoustic-based solutions in terms of portability, accuracy, usability, robustness, scalability, and responsiveness. Our extensive experiments with a prototype of${{\sf S \text{-}UbiTap}}$show that we can support sub-centimeter accuracy on various types of surfaces with minor user calibration effort. In addition, the accuracy is maintained even when the size of the touch input space increases. In our experience with real-world users,${{\sf S \text{-}UbiTap}}$significantly improves usability and robustness, thus enabling the emergence of more exciting applications. Anish Byanjankar, Yunxin Liu 0001, Yuanchao Shu, Insik Shin, Myeongwon Choi, Hyosu Kim |
IEEE Trans. Mob. Comput. | 2 |
| 2023 | LEAP: TrustZone Based Developer-Friendly TEE for Intelligent Mobile AppsabstractARM TrustZone is widely deployed on commercial-off-the-shelf mobile devices for secure execution. However, many Apps cannot enjoy this feature because it brings many constraints to App developers. Previous works have been proposed to build a secure execution environment for developers on top of TrustZone. Unfortunately, these works are still not a fully-fledged solution for mobile Apps, especially for the emerging intelligent Apps. To this end, we propose LEAP, which is a lightweight developer-friendly TEE solution for mobile Apps. LEAP enables isolated codes to execute in parallel and access peripheral (e.g., mobile GPUs) with ease, flexibly manages system resources upon different workloads, and offers the auto DevOps tool to help developers prepare the codes running on it. We implement the LEAP prototype on the off-the-shelf ARM platform and conduct extensive experiments on it. The experimental results show that Apps can be adapted to run with LEAP easily and efficiently. Compared to the state-of-the-art work along this research line, LEAP can achieve an average 3.57× speedup in supporting intelligent Apps using mobile GPU acceleration. Lizhi Sun, Shuocheng Wang, Hao Wu 0067, Yuhang Gong, Fengyuan Xu, Yunxin Liu 0001, Sheng Zhong 0002 |
IEEE Trans. Mob. Comput. | 6 |
| 2023 | MVPose: Realtime Multi-Person Pose Estimation Using Motion Vector on Mobile DevicesabstractWe present MVPose, a novel system designed to enable real-time multi-person pose estimation (PE) on commodity mobile devices, which consists of three novel techniques. First, MVPose takes a motion-vector-based approach to fast and accurately track the human keypoints across consecutive frames, rather than running expensive human-detection model and pose-estimation model for every frame. Second, MVPose designs a mobile-friendly PE model that uses lightweight feature extractors and multi-stage network to significantly reduce the latency of pose estimation without compromising the model accuracy. Third, MVPose leverages the heterogeneous computing resources of both CPU and GPU to execute the pose estimation model for multiple persons in parallel, which further reduces the total latency. We present extensive experiments to evaluate the effectiveness of the proposed tecniques by implemented the MVPose on five off-the-shelf commercial smartphones. Evaluation results show that MVPose achieves over30frames per second PE with4persons per frame, which significantly outperforms the state-of-the-art baseline, with a speedup of up to5.7×and3.8×in latency on CPU and GPU, respectively. Compared with baseline, MVPose achieves an improvement of10.1%in multi-person PE accuracy. Furthermore, MVPose achieves up to74.3%and57.6%energy-per-frame saving on average in comparison with the baseline on mobile CPU and GPU, respectively. Yunxin Liu 0001, Ju Ren 0001, Xiaohui Xu, Fucheng Jia, Yaoxue Zhang |
IEEE Trans. Mob. Comput. | 4 |
| 2022 | Nesting Forward Automatic Differentiation for Memory-Efficient Deep Neural Network TrainingabstractAn activation function is an element-wise mathematical function and plays a crucial role in deep neural networks (DNN). Many novel and sophisticated activation functions have been proposed to improve the DNN accuracy but also consume massive memory in the training process with back-propagation. In this study, we propose the nested forward automatic differentiation (Forward-AD), specifically for the element-wise activation function for memory-efficient DNN training. We deploy nested Forward-AD in two widely-used deep learning frameworks, TensorFlow and PyTorch, which support the static and dynamic computation graph, respectively. Our evaluation shows that nested Forward-AD reduces the memory footprint by up to 1.97× than the baseline model and outperforms the recomputation by 20% under the same memory reduction ratio. Cong Guo 0003, Yuxian Qiu, Jingwen Leng, Chen Zhang 0001, Quanlu Zhang, Yunxin Liu 0001, Fan Yang 0024, Minyi Guo |
ICCD | 7 |
| 2022 | SQuant: On-the-Fly Data-Free Quantization via Diagonal Hessian Approximation
Cong Guo 0003, Yuxian Qiu, Jingwen Leng, Xiaotian Gao, Chen Zhang 0001, Yunxin Liu 0001, Fan Yang 0024, Yuhao Zhu 0001, Minyi Guo |
ICLR | 6 |
| 2022 | Representational Continuity for Unsupervised Continual Learning
Divyam Madaan, Jaehong Yoon, Yuanchun Li 0003, Yunxin Liu 0001, Sung Ju Hwang |
ICLR | 4 |
| 2022 | ReMoS: Reducing Defect Inheritance in Transfer Learning via Relevant Model SlicingabstractTransfer learning is a popular software reuse technique in the deep learning community that enables developers to build custom models (students) based on sophisticated pretrained models (teachers). However, like vulnerability inheritance in traditional software reuse, some defects in the teacher model may also be inherited by students, such as well-known adversarial vulnerabilities and backdoors. Reducing such defects is challenging since the student is unaware of how the teacher is trained and/or attacked. In this paper, we propose ReMoS, a relevant model slicing technique to reduce defect inheritance during transfer learning while retaining useful knowledge from the teacher model. Specifically, ReMoS computes a model slice (a subset of model weights) that is relevant to the student task based on the neuron coverage information obtained by profiling the teacher model on the student task. Only the relevant slice is used to finetune the student model, while the irrelevant weights are retrained from scratch to minimize the risk of inheriting defects. Our experiments on seven DNN defects, four DNN models, and eight datasets demonstrate that ReMoS can reduce inherited defects effectively (by 63% to 86% for CV tasks and by 40% to 61% for NLP tasks) and efficiently with minimal sacrifice of accuracy (3% on average). Ziqi Zhang 0017, Yuanchun Li 0003, Jindong Wang 0001, Ding Li 0001, Yao Guo 0001, Xiangqun Chen, Yunxin Liu 0001 |
ICSE | 8 |
| 2022 | ANT: Exploiting Adaptive Numerical Data Type for Low-bit Deep Neural Network QuantizationabstractQuantization is a technique to reduce the computation and memory cost of DNN models, which are getting increasingly large. Existing quantization solutions use fixed-point integer or floating-point types, which have limited benefits, as both require more bits to maintain the accuracy of original models. On the other hand, variable-length quantization uses low-bit quantization for normal values and high-precision for a fraction of outlier values. Even though this line of work brings algorithmic benefits, it also introduces significant hardware overheads due to variable-length encoding and decoding.In this work, we propose a fixed-length a daptive n umerical data t ype called ANT to achieve low-bit quantization with tiny hardware overheads. Our data type ANT leverages two key innovations to exploit the intra-tensor and inter-tensor adaptive opportunities in DNN models. First, we propose a particular data type, flint, that combines the advantages of float and int for adapting to the importance of different values within a tensor. Second, we propose an adaptive framework that selects the best type for each tensor according to its distribution characteristics. We design a unified processing element architecture for ANT and show its ease of integration with existing DNN accelerators. Our design results in $2.8\times $ speedup and $2.5\times $ energy efficiency improvement over the state-of-the-art quantization accelerators. Cong Guo 0003, Chen Zhang 0001, Jingwen Leng, Zihan Liu 0002, Fan Yang 0024, Yunxin Liu 0001, Minyi Guo, Yuhao Zhu 0001 |
MICRO | 6 |
| 2022 | Romou: rapidly generate high-performance tensor kernels for mobile GPUsabstractMobile GPU, as a ubiquitous and powerful accelerator, plays an important role in accelerating on-device DNN (Deep Neural Network) inference. The frequent-upgrade and diversity of mobile GPUs require automatic kernel generation to empower fast DNN deployment. However, current generated kernels have poor performance. Rendong Liang, Ting Cao 0003, Jicheng Wen, Manni Wang, Yang Wang 0053, Jianhua Zou, Yunxin Liu 0001 |
MobiCom | 7 |
| 2022 | MobiDepth: real-time depth estimation using on-device dual camerasabstractReal-time depth estimation is critical for the increasingly popular augmented reality and virtual reality applications on mobile devices. Yet existing solutions are insufficient as they require expensive depth sensors or motion of the device, or have a high latency. We propose MobiDepth, a real-time depth estimation system using the widely-available on-device dual cameras. While binocular depth estimation is a mature technique, it is challenging to realize the technique on commodity mobile devices due to the different focal lengths and unsynchronized frame flows of the on-device dual cameras and the heavy stereo-matching algorithm. Ju Ren 0001, Bangwen He, Ting Cao 0003, Yuanchun Li 0003, Yaoxue Zhang, Yunxin Liu 0001 |
MobiCom | 9 |
| 2022 | CoDL: efficient CPU-GPU co-execution for deep learning inference on mobile devicesabstractConcurrent inference execution on heterogeneous processors is critical to improve the performance of increasingly heavy deep learning (DL) models. However, available inference frameworks can only use one processor at a time, or hardly achieve speedup by concurrent execution compared to using one processor. This is due to the challenges to 1) reduce data sharing overhead, and 2) properly partition each operator between processors. Fucheng Jia, Ting Cao 0003, Shiqi Jiang 0002, Yunxin Liu 0001, Ju Ren 0001, Yaoxue Zhang |
MobiSys | 5 |
| 2022 | FedBalancer: data and pace control for efficient federated learning on heterogeneous clientsabstractFederated Learning (FL) trains a machine learning model on distributed clients without exposing individual data. Unlike centralized training that is usually based on carefully-organized data, FL deals with on-device data that are often unfiltered and imbalanced. As a result, conventional FL training protocol that treats all data equally leads to a waste of local computational resources and slows down the global learning process. To this end, we propose FedBalancer, a systematic FL framework that actively selects clients' training samples. Our sample selection strategy prioritizes more "informative" data while respecting privacy and computational capabilities of clients. To better utilize the sample selection to speed up global training, we further introduce an adaptive deadline control scheme that predicts the optimal deadline for each round with varying client training data. Compared with existing FL algorithms with deadline configuration methods, our evaluation on five datasets from three different domains shows that FedBalancer improves the time-to-accuracy performance by 1.20~4.48× while improving the model accuracy by 1.1~5.0%. We also show that FedBalancer is readily applicable to other FL approaches by demonstrating that FedBalancer improves the convergence speed and accuracy when operating jointly with three different FL algorithms. Jaemin Shin 0005, Yuanchun Li 0003, Yunxin Liu 0001, Sung-Ju Lee 0001 |
MobiSys | 3 |
| 2022 | Melon: breaking the memory wall for resource-efficient on-device machine learningabstractOn-device learning is a promising technique for emerging privacy-preserving machine learning paradigms. However, through quantitative experiments, we find that commodity mobile devices cannot well support state-of-the-art DNN training with a large enough batch size, due to the limited local memory capacity. To fill the gap, we propose Melon, a memory-friendly on-device learning framework that enables the training tasks with large batch size beyond the physical memory capacity. Melon judiciously retrofits existing memory saving techniques to fit into resource-constrained mobile devices, i.e., recomputation and micro-batch. Melon further incorporates novel techniques to deal with the high memory fragmentation and memory adaptation. We implement and evaluate Melon with various typical DNN models on commodity mobile devices. The results show that Melon can achieve up to 4.33× larger batch size under the same memory budget. Given the same batch size, Melon achieves 1.89× on average (up to 4.01×) higher training throughput, and saves up to 49.43% energy compared to competitive alternatives. Furthermore, Melon reduces 78.59% computation on average in terms of memory budget adaptation. Qipeng Wang 0001, Mengwei Xu 0001, Chao Jin 0007, Xinran Dong, Jinliang Yuan, Xin Jin 0008, Gang Huang 0001, Yunxin Liu 0001, Xuanzhe Liu |
MobiSys | 8 |
| 2022 | TailorFL: Dual-Personalized Federated Learning under System and Data HeterogeneityabstractFederated learning (FL) enables distributed mobile devices to collaboratively learn a shared model without exposing their raw data. However, heterogeneous devices usually have limited and different available resources, i.e., system heterogeneity, for model training and communicating, while the diverse data distribution among devices, i.e., data heterogeneity, may result in significant performance degradation. In this paper, we propose TailorFL, a dual-personalized FL framework, which tailors a submodel for each device with personalized structure for training and personalized parameters for local inference. To achieve this, we first excavate the personalization principle for data heterogeneous FL via in-depth empirical studies, and based on which, we propose a resource-aware and data-directed pruning strategy that makes each device's submodel structure match its resource capability and correlate with its local data distribution. To aggregate the submodels while preserving their dual personalization properties, we design a scaling-based aggregation strategy that scales parameters with the pruning rate of submodels and aggregates the overlapped parameters. Moreover, to further promote beneficial and restrain detrimental collaborations among devices, we propose a server-assisted model-tuning mechanism, which dynamically tunes device's submodel structure at the server side with the global view of device's data distribution similarities. Extensive experiments demonstrate that compared to the status quo approaches, TailorFL achieves an average of 22% increase in inference accuracy, and reduces the memory, computation, and communication costs for model training simultaneously. Yongheng Deng, Weining Chen, Ju Ren 0001, Feng Lyu 0001, Yang Liu 0165, Yunxin Liu 0001, Yaoxue Zhang |
SenSys | 6 |
| 2022 | Hyperion: A Generic and Distributed Mobile Offloading Framework on OpenCLabstractDespite the significant development of mobile device SoCs, they are still inefficient in computing computation-intensive workloads, such as high-resolution image processing and AR/VR applications. Offloading offers a promising way to leverage cloud or edge servers for acceleration, but existing offloading is limited to specific tasks or specific hardware/software platforms, resulting in significant engineering overhead. To address this problem, we focus on the underlying layer of these applications (i.e., OpenCL) and propose Hyperion, a generic and distributed mobile offloading framework built on OpenCL. To achieve high-performance distributed execution for Hyperion, we first take a deep insight into the OpenCL data structures and design regularity-aware kernel analyzer to analyze the data dependency of work-groups and identify the essential data to offload. Then, context-aware execution time predictor is proposed to estimate the computing time of a given partitioned kernel workload that is highly impacted by many runtime factors. These techniques are integrated into pipeline-enabled and network-adaptive scheduler to make scheduling decisions, which coordinates the kernel partition and workload scheduling to form pipeline processing between data transmission and distributed execution with flexible adaptability to network dynamics. Extensive experimental results demonstrate that Hyperion achieves superior performance with an average 3.80× speedup compared with the best baseline and flexible adaptation to dynamic network conditions and available computing resources. Ziyan Fu 0001, Ju Ren 0001, Yunxin Liu 0001, Ting Cao 0003, Yue-Zhi Zhou, Yaoxue Zhang |
SenSys | 3 |
| 2022 | WheelLoc: Practical and Accurate Localization for Wheeled Mobile Targets via Integrated Sensing and CommunicationabstractPractical and accurate localization systems are important to mobile targets that enable promising services such as navigation and augmented reality. With the proliferation of WiFi, existing WiFi-based localization systems have leveraged RSSI, fingerprints, landmarks, time of arrival, or angle of arrival to locate targets, while no related work pays attention to mobile targets themselves. For wheel-driven mobile targets, such as vehicles, bikes, and wheeled robots, we design and implement WheelLoc, a novel WiFi-based localization system leveraging the rotation of wheels. The specially designed WheelLoc hardware is cost-effective and self-powered with the composition of three commercial antennas and a solar cell, which is also easy to be installed on wheels. A hybrid WheelLoc algorithm is further proposed to realize accurate localization in diverse environments, whether the wheel of targets is static or mobile, indoor or outdoor, on flat or bumpy ground. The movements of individual antennas are exploited to emulate linear, cycloid, and circular antenna arrays using a new formulation of Synthetic Aperture Radar (SAR). Extensive experiments are conducted on bikes in the real world. Performance results demonstrate that WheelLoc does not require any user interaction, yet achieves comparable accuracy with the state-of-the-art localization systems using WiFi. Linghe Kong, Yunxin Liu 0001, Le Zheng, Meikang Qiu, Guihai Chen |
IEEE J. Sel. Areas Commun. | 3 |
| 2022 | A Cloud-Edge Collaboration Framework for Cognitive ServiceabstractMobile applications can leverage high-quality deep learning models such as convolutional neural networks and deep neural networks to provide high-performance cognitive services. Prior work on deep learning models-based mobile applications in a cloud-edge computing environment focuses on performing lightweight data pre-processing tasks on edge servers for cloud-hosted cognitive servers. These approaches have two major limitations. First, it is uneasy for the mobile applications to assure satisfactory user experience in terms of network communication delay, because the intermediary edge servers are used only to pre-process data (e.g., images and videos) and the cloud servers are used to complete the tasks. Second, these approaches assume the pre-trained deep learning models deployed on cloud servers are static, and will not attempt to automatically upgrade in a context-aware manner. In this article, we propose a cloud-edge collaboration framework that facilitates delivering cognitive services with long-lasting, fast response, and high accuracy properties. We fist deploy a shallow model (i.e., EdgeCNN) on the edge server and a deep model (i.e., CloudCNN) on the cloud server. EdgeCNN can provide durable and rapid response cognitive services, because edge servers not only provide computing resources for mobile applications, but also close to users. Then, we enable CloudCNN to assist in training EdgeCNN to improve the performance of the latter. Thus, EdgeCNN also provides high-accuracy cognitive services. Furthermore, because users may continue to upload data to edge servers in real-world scenarios, we propose to use the ongoing assistance of CloudCNN to further improve the accuracy of the shallow model. Experimental results show that EdgeCNN can reduce the average response time of cognitive services by up to 55.08 percent and improve accuracy by up to 26.70 percent. Chuntao Ding, Ao Zhou 0001, Yunxin Liu 0001, Rong Chang 0001, Ching-Hsien Hsu, Shangguang Wang |
IEEE Trans. Cloud Comput. | 3 |
| 2022 | Model Protection: Real-Time Privacy-Preserving Inference Service for Model Privacy at the EdgeabstractMajor cloud service providers with well-equipped infrastructure, experienced machine learning (ML) expertise, and enriched training datasets are building ML-as-a-Service (MLaaS) systems, in which clients can query ML-based prediction services with their data. Instead of moving private data to the cloud, in this work, we design, implement, and evaluate a novel secure ML system to enable MLaaS on edge devices. To protect the proprietary ML models on edge devices from revealing to the clients while maintaining a real-time inference is challenging. Existing privacy-preserving ML techniques can hardly satisfy real-time requirements. In our solution, we employ a secure enclave (e.g., SGX) to offer security and provide better efficiency than cryptographic techniques. However, the enclave alone cannot achieve real-time capability due to its limited capacity. We observe that the ML model imposes a severe accuracy degradation when adding noise to a few model weights. Based on this, we design a suite of novel solutions to optimize the performance of secure enclave-based inference service at the edge by enclosing only$1\%$computation within secure enclaves. Our work can achieve up to a$7.8\times$increase in efficiency and a$27\times$reduction in memory usage compared to the state-of-the-art. Jiahui Hou, Huiqi Liu, Yunxin Liu 0001, Yu Wang 0003, Peng-Jun Wan, Xiang-Yang Li 0001 |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2022 | Characterizing Embedded Web Browsing in Mobile AppsabstractModern mobile OSes support to display Web pages in the native apps, which we call embedded Web pages. In this paper, we conduct, to the best of our knowledge, the first measurement study on browsing embedded Web pages on Android. Our study on 22,521 popular Android apps shows that 57.9% and 73.8% of apps embed Web pages on two popular app markets: Google Play and Wandoujia, respectively. To analyze the embedded Web browsing performance at scale, we design and implement EWProfiler, a tool that can automatically search for embedded Web pages inside apps, trigger page loads, and retrieve performance metrics. Based on 445 embedded Web pages obtained by EWProfiler in 99 popular apps from the two app markets, we investigate the characteristics and performance of embedded Web pages, and find that embedded Web pages significantly impede the app user experience. To optimize the performance of embedded Web browsing, we investigate the effectiveness of three techniques, i.e., separating the browser kernel to a different process, loading pages from local storage, and pre-rendering. We believe that our findings could draw attentions to Web developers, browser vendors, app developers, and mobile OS vendors together towards better performance of embedded Web browsing. Deyu Tian, Yun Ma 0002, Aruna Balasubramanian, Yunxin Liu 0001, Gang Huang 0001, Xuanzhe Liu |
IEEE Trans. Mob. Comput. | 4 |
| 2021 | DeepPayload: Black-box Backdoor Attack on Deep Learning Models through Neural Payload InjectionabstractDeep learning models are increasingly used in mobile applications as critical components. Unlike the program bytecode whose vulnerabilities and threats have been widely-discussed, whether and how the deep learning models deployed in the applications can be compromised are not well-understood since Neural Networks are usually viewed as a black box. In this paper, we introduce a highly practical backdoor attack achieved with a set of reverse-engineering techniques over compiled deep learning models. The core of the attack is a neural conditional branch constructed with a trigger detector and several operators and injected into the victim model as a malicious payload. The attack is effective as the conditional logic can be flexibly customized by the attacker, and scalable as it does not require any prior knowledge from the original model. We evaluated the attack effectiveness using 5 state-of-the-art deep learning models and real-world samples collected from 30 users. The results demonstrated that the injected backdoor can be triggered with a success rate of 93.5%, while only brought less than 2ms latency overhead and no more than 1.4% accuracy decrease. We further conducted an empirical study on real-world mobile deep learning apps collected from Google Play. We found 54 apps that were vulnerable to our attack, including popular and security-critical ones. The results call for the awareness of deep learning application developers and auditors to enhance the protection of deployed models. Yuanchun Li 0003, Jiayi Hua, Haoyu Wang 0001, Chunyang Chen 0001, Yunxin Liu 0001 |
ICSE | 5 |
| 2021 | Dual-side Sparse Tensor CoreabstractLeveraging sparsity in deep neural network (DNN) models is promising for accelerating model inference. Yet existing GPUs can only leverage the sparsity from weights but not activations, which are dynamic, unpredictable, and hence challenging to exploit. In this work, we propose a novel architecture to efficiently harness the dual-side sparsity (i.e., weight and activation sparsity). We take a systematic approach to understand the (dis)advantages of previous sparsity-related architectures and propose a novel, unexplored paradigm that combines outer-product computation primitive and bitmap-based encoding format. We demonstrate the feasibility of our design with minimal changes to the existing production-scale inner-product-based Tensor Core. We propose a set of novel ISA extensions and co-design the matrix-matrix multiplication and convolution algorithms, which are the two dominant computation patterns in today’s DNN models, to exploit our new dual-side sparse Tensor Core. Our evaluation shows that our design can fully unleash the dual-side DNN sparsity and improve the performance by up to one order of magnitude with small hardware overhead. Yang Wang 0053, Chen Zhang 0001, Cong Guo 0003, Yunxin Liu 0001, Jingwen Leng |
ISCA | 5 |
| 2021 | ModelDiff: testing-based DNN similarity comparison for model reuse detectionabstractThe knowledge of a deep learning model may be transferred to a student model, leading to intellectual property infringement or vulnerability propagation. Detecting such knowledge reuse is nontrivial because the suspect models may not be white-box accessible and/or may serve different tasks. In this paper, we propose ModelDiff, a testing-based approach to deep learning model similarity comparison. Instead of directly comparing the weights, activations, or outputs of two models, we compare their behavioral patterns on the same set of test inputs. Specifically, the behavioral pattern of a model is represented as a decision distance vector (DDV), in which each element is the distance between the model's reactions to a pair of inputs. The knowledge similarity between two models is measured with the cosine similarity between their DDVs. To evaluate ModelDiff, we created a benchmark that contains 144 pairs of models that cover most popular model reuse methods, including transfer learning, model compression, and model stealing. Our method achieved 91.7% correctness on the benchmark, which demonstrates the effectiveness of using ModelDiff for model reuse detection. A study on mobile deep learning apps has shown the feasibility of ModelDiff on real-world models. Yuanchun Li 0003, Ziqi Zhang 0017, Yunxin Liu 0001 |
ISSTA | 5 |
| 2021 | Boosting Mobile CNN Inference through Semantic MemoryabstractHuman brains are known to be capable of speeding up visual recognition of repeatedly presented objects through faster memory encoding and accessing procedures on activated neurons. For the first time, we borrow and distill such a capability into a semantic memory design, namely SMTM, to improve on-device CNN inference. SMTM employs a hierarchical memory architecture to leverage the long-tail distribution of objects of interest, and further incorporates several novel techniques to put it into effects: (1) it encodes high-dimensional feature maps into low-dimensional, semantic vectors for low-cost yet accurate cache and lookup; (2) it uses a novel metric in determining the exit timing considering different layers' inherent characteristics; (3) it adaptively adjusts the cache size and semantic vectors to fit the scene dynamics. SMTM is prototyped on commodity CNN engine and runs on both mobile CPU and GPU. Extensive experiments on large-scale datasets and models show that SMTM can significantly speed up the model inference over standard approach (up to 2×) and prior cache designs (up to 1.5x), with acceptable accuracy loss. Chen Zhang 0001, Shihao Han, Li Lyna Zhang, Baoqun Yin, Yunxin Liu 0001, Mengwei Xu 0001 |
ACM Multimedia | 6 |
| 2021 | Flexible high-resolution object detection on edge devices with tunable latencyabstractObject detection is a fundamental building block of video analytics applications. While Neural Networks (NNs)-based object detection models have shown excellent accuracy on benchmark datasets, they are not well positioned for high-resolution images inference on resource-constrained edge devices. Common approaches, including down-sampling inputs and scaling up neural networks, fall short of adapting to video content changes and various latency requirements. This paper presents Remix, a flexible framework for high-resolution object detection on edge devices. Remix takes as input a latency budget, and come up with an image partition and model execution plan which runs off-the-shelf neural networks on non-uniformly partitioned image blocks. As a result, it maximizes the overall detection accuracy by allocating various amount of compute power onto different areas of an image. We evaluate Remix on public dataset as well as real-world videos collected by ourselves. Experimental results show that Remix can either improve the detection accuracy by 18%-120% for a given latency budget, or achieve up to 8.1× inference speedup with accuracy on par with the state-of-the-art NNs. Shiqi Jiang 0002, Yuanchun Li 0003, Yuanchao Shu, Yunxin Liu 0001 |
MobiCom | 5 |
| 2021 | AsyMo: scalable and efficient deep-learning inference on asymmetric mobile CPUsabstractOn-device deep learning (DL) inference has attracted vast interest. Mobile CPUs are the most common hardware for on-device inference and many inference frameworks have been developed for them. Yet, due to the hardware complexity, DL inference on mobile CPUs suffers from two common issues: the poor performance scalability on the asymmetric multiprocessor, and energy inefficiency. Manni Wang, Shaohua Ding, Ting Cao 0003, Yunxin Liu 0001, Fengyuan Xu |
MobiCom | 4 |
| 2021 | PECAM: privacy-enhanced video streaming and analytics via securely-reversible transformationabstractAs Video Streaming and Analytics (VSA) systems become increasingly popular, serious privacy concerns have risen on exposing too much unnecessary private information to the VSA providers. Yet, it is challenging to protect privacy while still preserving desired VSA features, i.e., effective analytics, forensic support, resource efficiency, and real-time execution. In this paper, we present a VSA privacy enhancement system (PECAM), which addresses the above challenge with no change in the VSA back-end. PECAM leverages a novel Generative Adversarial Network to perform the privacy-enhanced securely-reversible video transformation. PECAM also incorporates a couple of system optimizations into its VSA workflow to reduce network bandwidth usage and enable real-time processing on cameras. We implement our PECAM prototype on commodity hardware and evaluate its performance via both security study and extensive experiments. Results demonstrate that PECAM can effectively enhance the visual privacy of VSA in the presence of an adversary, and its transformed videos, when taken as input for various VSA back-end tasks, maintain a 96% accuracy of corresponding original videos. Additionally, it performs 12.3× and 1.8× better than baseline methods in terms of the computing cost and network bandwidth usage, respectively. Hao Wu 0067, Xuejin Tian, Minghao Li 0003, Yunxin Liu 0001, Ganesh Ananthanarayanan, Fengyuan Xu, Sheng Zhong 0002 |
MobiCom | 4 |
| 2021 | Elf: accelerate high-resolution mobile deep vision with content-aware parallel offloadingabstractAs mobile devices continuously generate streams of images and videos, a new class of mobile deep vision applications are rapidly emerging, which usually involve running deep neural networks on these multimedia data in real-time. To support such applications, having mobile devices offload the computation, especially the neural network inference, to edge clouds has proved effective. Existing solutions often assume there exists a dedicated and powerful server, to which the entire inference can be offloaded. In reality, however, we may not be able to find such a server but need to make do with less powerful ones. To address these more practical situations, we propose to partition the video frame and offload the partial inference tasks to multiple servers for parallel processing. This paper presents the design of Elf, a framework to accelerate the mobile deep vision applications with any server provisioning through the parallel offloading. Elf employs a recurrent region proposal prediction algorithm, a region proposal centric frame partitioning, and a resource-aware multi-offloading scheme. We implement and evaluate Elf upon Linux and Android platforms using four commercial mobile devices and three deep vision applications with ten state-of-the-art models. The comprehensive experiments show that Elf can speed up the applications by 4.85× with saving bandwidth usage by 52.6%, while with <1% application accuracy sacrifice. Wuyang Zhang, Zhezhi He, Zhenhua Jia, Yunxin Liu 0001, Marco Gruteser, Dipankar Raychaudhuri, Yanyong Zhang |
MobiCom | 5 |
| 2021 | nn-Meter: towards accurate latency prediction of deep-learning model inference on diverse edge devicesabstractWith the recent trend of on-device deep learning, inference latency has become a crucial metric in running Deep Neural Network (DNN) models on various mobile and edge devices. To this end, latency prediction of DNN model inference is highly desirable for many tasks where measuring the latency on real devices is infeasible or too costly, such as searching for efficient DNN models with latency constraints from a huge model-design space. Yet it is very challenging and existing approaches fail to achieve a high accuracy of prediction, due to the varying model-inference latency caused by the runtime optimizations on diverse edge devices. Li Lyna Zhang, Shihao Han, Jianyu Wei, Ningxin Zheng, Ting Cao 0003, Yuqing Yang 0001, Yunxin Liu 0001 |
MobiSys | 7 |
| 2021 | TaintStream: fine-grained taint tracking for big data platforms through dynamic code translationabstractBig data has become valuable property for enterprises and enabled various intelligent applications. Today, it is common to host data in big data platforms (e.g., Spark), where developers can submit scripts to process the original and intermediate data tables. Meanwhile, it is highly desirable to manage the data to comply with various privacy requirements. To enable flexible and automated privacy policy enforcement, we propose TaintStream, a fine-grained taint tracking framework for Spark-like big data platforms. TaintStream works by automatically injecting taint tracking logic into the data processing scripts, and the injected scripts are dynamically translated to maintain a taint tag for each cell during execution. The dynamic translation rules are carefully designed to guarantee non-interference in the original data operation. By defining different semantics of taint tags, TaintStream can enable various data management applications such as access control, data retention, and user data erasure. Our experiments on a self-crafted benchmarksuite show that TaintStream is able to achieve accurate cell-level taint tracking with a precision of 93.0% and less than 15% overhead. We also demonstrate the usefulness of TaintStream through several real-world use cases of privacy policy enforcement. Chengxu Yang, Yuanchun Li 0003, Mengwei Xu 0001, Zhenpeng Chen 0001, Yunxin Liu 0001, Gang Huang 0001, Xuanzhe Liu |
ESEC/SIGSOFT FSE | 5 |
| 2021 | Video Analytics with Zero-streaming Cameras
Mengwei Xu 0001, Tiantu Xu, Yunxin Liu 0001, Felix Xiaozhu Lin |
USENIX ATC | 3 |
| 2021 | Characterizing Impacts of Heterogeneity in Federated Learning upon Large-Scale Smartphone DataabstractFederated learning (FL) is an emerging, privacy-preserving machine learning paradigm, drawing tremendous attention in both academia and industry. A unique characteristic of FL is heterogeneity, which resides in the various hardware specifications and dynamic states across the participating devices. Theoretically, heterogeneity can exert a huge influence on the FL training process, e.g., causing a device unavailable for training or unable to upload its model updates. Unfortunately, these impacts have never been systematically studied and quantified in existing FL literature. Chengxu Yang, Qipeng Wang 0001, Mengwei Xu 0001, Zhenpeng Chen 0001, Kaigui Bian, Yunxin Liu 0001, Xuanzhe Liu |
WWW | 6 |
| 2021 | S2Net: Preserving Privacy in Smart Home RoutersabstractAt present, wireless home routers are becoming increasingly smart. While these smart routers provide rich functionalities to users, they also raise security concerns. Although the existing end-to-end encryption techniques can be applied to protect personal data, such rich functionalities become unavailable due to the encrypted payloads. On the other hand, if the smart home routers are allowed to process and store the personal data of users, once compromised, the users' sensitive data will be exposed. As a consequence, users face a difficult trade-off between the benefits of the rich functionalities and potential privacy risks. To deal with this dilemma, we propose a novel system named Secure and Smart Network (S2Net) for home routers. For S2Net, we propose a secure OS that can distinguish and manage multiple sessions belonging to different users. The secure OS and all the router applications are placed in the secure world using the ARM TrustZone technology. In S2Net, we also confine the router applications in sandboxes provided by the proposed secure OS to prevent data leakage. As a result, S2Net can provide rich functionalities for users while preserving strong privacy for home routers. In addition, we develop a crypto-worker model that provides an abstraction layer of cryptographic tasks performed by a heterogeneous multi-core system. The other important role of crypto-worker is to parallelize the computations in order to resolve the high computation cost of cryptographic functions. We report the system design of S2Net and the details of our implementation. Experimental results with benchmarks and real applications demonstrate that our implementation is capable of achieving high performance in terms of throughput while mitigating the overhead of S2Net design. SeungSeob Lee, Kun Tan 0002, Yunxin Liu 0001, Yong Cui 0001 |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2021 | EPASS360: QoE-Aware 360-Degree Video Streaming Over Mobile DevicesabstractThe 360-degree video streaming system delivers a monocular panoramic video surrounding the user, and the user can change the viewing direction of mobile devices to see different parts of the video through the “viewport”. Due to the limited network bandwidth, playbacks of high-resolution 360-degree videos often suffer from rebuffering, while too much bandwidth is wasted in delivering those out-of-viewport parts that the user never watches. In this article, we present an Ensemble Prediction and Allocation based Streaming System, named as EPASS360, for delivering high Quality of Experience (QoE) 360-degree videos. The prediction model takes advantages of ensemble learning, providing high accuracy on the prediction of viewports. The allocation model divides a video into tiles, and allocates high resolution to tiles where a user's viewpoint may appear in the future by solving the QoE-aware optimization problem. Trace-driven emulation on real-world datasets shows that EPASS360 enhances the QoE in various scenarios compared to state-of-the-art streaming approaches. Experiments on the head-mounted device and the hand-held device over real-world Internet confirm the high user experience of EPASS360. Yuanxing Zhang, Yushuo Guan, Kaigui Bian, Yunxin Liu 0001, Hu Tuo, Lingyang Song, Xiaoming Li 0001 |
IEEE Trans. Mob. Comput. | 4 |
| 2021 | Operating Systems for Resource-adaptive Intelligent Software: Challenges and OpportunitiesabstractThe past decades witnessed the fast and wide deployment of Internet. The Internet has bred the ubiquitous computing environment that is spanning the cloud, edge, mobile devices, and IoT. Software running over such a ubiquitous computing environment environment is eating the world. A recently emerging trend of Internet-based software systems is “ resource adaptive ,” i.e., software systems should be robust and intelligent enough to the changes of heterogeneous resources, both physical and logical, provided by their running environment. To keep pace of such a trend, we argue that some considerations should be taken into account for the future operating system design and implementation. From the structural perspective, rather than the “monolithic OS” that manages the aggregated resources on the single machine, the OS should be dynamically composed over the distributed resources and flexibly adapt to the resource and environment changes. Meanwhile, the OS should leverage advanced machine/deep learning techniques to derive configurations and policies and automatically learn to tune itself and schedule resources. This article envisions our recent thinking of the new OS abstraction, namely, ServiceOS , for future resource-adaptive intelligent software systems. The idea of ServiceOS is inspired by the delivery model of “ Software-as-a-Service ” that is supported by the Service-Oriented Architecture (SOA). The key principle of ServiceOS is based on resource disaggregation, resource provisioning as a service, and learning-based resource scheduling and allocation. The major goal of this article is not providing an immediately deployable OS. Instead, we aim to summarize the challenges and potentially promising opportunities and try to provide some practical implications for researchers and practitioners. Xuanzhe Liu, Shangguang Wang, Yun Ma 0002, Ying Zhang 0012, Qiaozhu Mei, Yunxin Liu 0001, Gang Huang 0001 |
ACM Trans. Internet Techn. | 6 |
| 2021 | Efficient Data Loader for Fast Sampling-Based GNN Training on Large GraphsabstractEmerging graph neural networks (GNNs) have extended the successes of deep learning techniques against datasets like images and texts to more complex graph-structured data. By leveraging GPU accelerators, existing frameworks combine mini-batch and sampling for effective and efficient model training on large graphs. However, this setup faces a scalability issue since loading rich vertex features from CPU to GPU through a limited bandwidth link usually dominates the training cycle. In this article, we propose PaGraph, a novel, efficient data loader that supports general and efficient sampling-based GNN training on single-server with multi-GPU. PaGraph significantly reduces the data loading time by exploiting available GPU resources to keep frequently-accessed graph data with a cache. It also embodies a lightweight yet effective caching policy that takes into account graph structural information and data access patterns of sampling-based GNN training simultaneously. Furthermore, to scale out on multiple GPUs, PaGraph develops a fast GNN-computation-aware partition algorithm to avoid cross-partition access during data-parallel training and achieves better cache efficiency. Finally, it overlaps data loading and GNN computation for further hiding loading costs. Evaluations on two representative GNN models, GCN and GraphSAGE, using two sampling methods, Neighbor and Layer-wise, show that PaGraph could eliminate the data loading time from the GNN training pipeline, and achieve up to 4.8× performance speedup over the state-of-the-art baselines. Together with preprocessing optimization, PaGraph further delivers up to 16.0× end-to-end speedup. Youhui Bai, Cheng Li 0001, Yufei Wu 0011, Youshan Miao, Yunxin Liu 0001, Yinlong Xu 0001 |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2020 | PaGraph: Scaling GNN training on large graphs via computation-aware cachingabstractEmerging graph neural networks (GNNs) have extended the successes of deep learning techniques against datasets like images and texts to more complex graph-structured data. By leveraging GPU accelerators, existing frameworks combine both mini-batch and sampling for effective and efficient model training on large graphs. However, this setup faces a scalability issue since loading rich vertices features from CPU to GPU through a limited bandwidth link usually dominates the training cycle. In this paper, we propose PaGraph, a system that supports general and efficient sampling-based GNN training on single-server with multi-GPU. PaGraph significantly reduces the data loading time by exploiting available GPU resources to keep frequently accessed graph data with a cache. It also embodies a lightweight yet effective caching policy that takes into account graph structural information and data access patterns of sampling-based GNN training simultaneously. Furthermore, to scale out on multiple GPUs, PaGraph develops a fast GNN-computation-aware partition algorithm to avoid cross-partition access during data parallel training and achieves better cache efficiency. Evaluations on two representative GNN models, GCN and GraphSAGE, show that PaGraph achieves up to 96.8% data loading time reductions and up to 4.8X performance speedup over the state-of-the-art baselines. Together with preprocessing optimization, PaGraph further delivers up to 16.0X end-to-end speedup. Cheng Li 0001, Youshan Miao, Yunxin Liu 0001, Yinlong Xu 0001 |
SoCC | 4 |
| 2020 | SCYLLA: QoE-aware Continuous Mobile Vision with FPGA-based Dynamic Deep Neural Network ReconfigurationabstractContinuous mobile vision is becoming increasingly important as it finds compelling applications which substantially improve our everyday life. However, meeting the requirements of quality of experience (QoE) diversity, energy efficiency and multi-tenancy simultaneously represents a significant challenge. In this paper, we present SCYLLA, an FPGA-based framework that enables QoE-aware continuous mobile vision with dynamic reconfiguration to effectively address this challenge. SCYLLA pre-generates a pool of FPGA design and DNN models, and dynamically applies the optimal software-hardware configuration to achieve the maximum overall performance on QoE for concurrent tasks. We implement SCYLLA on state-of-the-art FPGA platform and evaluate SCYLLA using drone-based traffic surveillance application on three datasets. Our evaluation shows that SCYLLA provides much better design flexibility and achieves superior QoE trade-offs than status-quo CPU-based solution that existing continuous mobile vision applications are built upon. Shuang Jiang, Zhiyao Ma, Chenren Xu, Mi Zhang 0002, Chen Zhang 0001, Yunxin Liu 0001 |
INFOCOM | 7 |
| 2020 | A query engine for zero-streaming camerasabstractLow-cost wireless cameras are growing rapidly. With the help of advanced machine learning models (e.g., CNNs), those videos exhibit high business and social values, e.g., for retailing planning [18], wildlife study [21], and traffic monitoring [19, 25]. However, with high compute need, traditional video analytics systems [14, 15, 26, 27] require all videos to be uploaded to a backend server, which stresses the scarce network bandwidth between cameras and servers. Mengwei Xu 0001, Tiantu Xu, Yunxin Liu 0001, Xuanzhe Liu, Gang Huang 0001, Felix Xiaozhu Lin |
MobiCom | 3 |
| 2020 | EMO: real-time emotion recognition from single-eye images for resource-constrained eyewear devicesabstractReal-time user emotion recognition is highly desirable for many applications on eyewear devices like smart glasses. However, it is very challenging to enable this capability on such devices due to tightly constrained image contents (only eye-area images available from the on-device eye-tracking camera) and computing resources of the embedded system. In this paper, we propose and develop a novel system called EMO that can recognize, on top of a resource-limited eyewear device, real-time emotions of the user who wears it. Unlike most existing solutions that require whole-face images to recognize emotions, EMO only utilizes the single-eye-area images captured by the eye-tracking camera of the eyewear. To achieve this, we design a customized deep-learning network to effectively extract emotional features from input single-eye images and a personalized feature classifier to accurately identify a user's emotions. EMO also exploits the temporal locality and feature similarity among consecutive video frames of the eye-tracking camera to further reduce the recognition latency and system resource usage. We implement EMO on two hardware platforms and conduct comprehensive experimental evaluations. Our results demonstrate that EMO can continuously recognize seven-type emotions at 12.8 frames per second with a mean accuracy of 72.2%, significantly outperforming the state-of-the-art approach, and consume much fewer system resources. Hao Wu 0067, Xuejin Tian, Edward Sun, Yunxin Liu 0001, Fengyuan Xu, Sheng Zhong 0002 |
MobiSys | 5 |
| 2020 | Approximate query service on autonomous IoT camerasabstractElf is a runtime for an energy-constrained camera to continuously summarize video scenes as approximate object counts. Elf's novelty centers on planning the camera's count actions under energy constraint. (1) Elf explores the rich action space spanned by the number of sample image frames and the choice of per-frame object counters; it unifies errors from both sources into one single bounded error. (2) To decide count actions at run time, Elf employs a learning-based planner, jointly optimizing for past and future videos without delaying result materialization. Tested with more than 1,000 hours of videos and under realistic energy constraints, Elf continuously generates object counts within only 11% of the true counts on average. Alongside the counts, Elf presents narrow errors shown to be bounded and up to 3.4X smaller than competitive baselines. At a higher level, Elf makes a case for advancing the geographic frontier of video analytics. Mengwei Xu 0001, Yunxin Liu 0001, Gang Huang 0001, Xuanzhe Liu, Felix Xiaozhu Lin |
MobiSys | 3 |
| 2020 | MobiPose: real-time multi-person pose estimation on mobile devicesabstractHuman pose estimation is a key technique for many vision-based mobile applications. Yet existing multi-person pose-estimation methods fail to achieve a satisfactory user experience on commodity mobile devices such as smartphones, due to their long model-inference latency. In this paper, we propose MobiPose, a system designed to enable real-time multi-person pose estimation on mobile devices through three novel techniques. First, MobiPose takes a motion-vector-based approach to fast locate the human proposals across consecutive frames by fine-grained tracking of joints of human body, rather than running the expensive human-detection model for every frame. Second, MobiPose designs a mobile-friendly model that uses lightweight multi-stage feature extractions to significantly reduce the latency of pose estimation without compromising the model accuracy. Third, MobiPose leverages the heterogeneous computing resources of both CPU and GPU to execute the pose estimation model for multiple persons in parallel, which further reduces the total latency. We have implemented the MobiPose system on off-the-shelf commercial smartphones and conducted comprehensive experiments to evaluate the effectiveness of the proposed techniques. Evaluation results show that MobiPose achieves over 20 frames per second pose estimation with 3 persons per frame, and significantly outperforms the state-of-the-art baseline, with a speedup of up to 4.5X and 2.8X in latency on CPU and GPU, respectively, and an improvement of 5.1% in pose-estimation model accuracy. Furthermore, MobiPose achieves up to 62.5% and 37.9% energy-per-frame saving on average in comparison with the baseline on mobile CPU and GPU, respectively. Xiaohui Xu, Fucheng Jia, Yunxin Liu 0001, Xuanzhe Liu, Ju Ren 0001, Yaoxue Zhang |
SenSys | 5 |
| 2020 | Dynamic slicing for deep neural networksabstractProgram slicing has been widely applied in a variety of software engineering tasks. However, existing program slicing techniques only deal with traditional programs that are constructed with instructions and variables, rather than neural networks that are composed of neurons and synapses. In this paper, we introduce NNSlicer, the first approach for slicing deep neural networks based on data-flow analysis. Our method understands the reaction of each neuron to an input based on the difference between its behavior activated by the input and the average behavior over the whole dataset. Then we quantify the neuron contributions to the slicing criterion by recursively backtracking from the output neurons, and calculate the slice as the neurons and the synapses with larger contributions. We demonstrate the usefulness and effectiveness of NNSlicer with three applications, including adversarial input detection, model pruning, and selective model protection. In all applications, NNSlicer significantly outperforms other baselines that do not rely on data flow analysis. Ziqi Zhang 0017, Yuanchun Li 0003, Yao Guo 0001, Xiangqun Chen, Yunxin Liu 0001 |
ESEC/SIGSOFT FSE | 5 |
| 2020 | Distributed fine-tuning of CNNs for image retrieval on multiple mobile devices
Gwangseon Jang, Jae-Gil Lee 0001, Yunxin Liu 0001 |
Pervasive Mob. Comput. | 4 |
| 2019 | SeerNet: Predicting Convolutional Neural Network Feature-Map Sparsity Through Low-Bit QuantizationabstractIn this paper we present a novel and general method to accelerate convolutional neural network (CNN) inference by taking advantage of feature map sparsity. We experimentally demonstrate that a highly quantized version of the original network is sufficient in predicting the output sparsity accurately, and verify that leveraging such sparsity in inference incurs negligible accuracy drop compared with the original network. To accelerate inference, for each convolution layer our approach first obtains a binary sparsity mask of the output feature maps by running inference on a quantized version of the original network layer, and then conducts a full-precision sparse convolution to find out the precise values of the non-zero outputs. Compared with existing work, our approach avoids the overhead of training additional auxiliary networks, while is still applicable to general CNN networks without being limited to certain application domains. Shijie Cao, Lingxiao Ma, Wencong Xiao, Chen Zhang 0001, Yunxin Liu 0001, Lanshun Nie, Zhi Yang 0001 |
CVPR | 5 |
| 2019 | Efficient and Effective Sparse LSTM on FPGA with Bank-Balanced SparsityabstractNeural networks based on Long Short-Term Memory (LSTM) are widely deployed in latency-sensitive language and speech applications. To speed up LSTM inference, previous research proposes weight pruning techniques to reduce computational cost. Unfortunately, irregular computation and memory accesses in unrestricted sparse LSTM limit the realizable parallelism, especially when implemented on FPGA. To address this issue, some researchers propose block-based sparsity patterns to increase the regularity of sparse weight matrices, but these approaches suffer from deteriorated prediction accuracy. This work presents Bank-Balanced Sparsity (BBS), a novel sparsity pattern that can maintain model accuracy at a high sparsity level while still enable an efficient FPGA implementation. BBS partitions each weight matrix row into banks for parallel computing, while adopts fine-grained pruning inside each bank to maintain model accuracy. We develop a 3-step software-hardware co-optimization approach to apply BBS in real FPGA hardware. First, we propose a bank-balanced pruning method to induce the BBS pattern on weight matrices. Then we introduce a decoding-free sparse matrix format, Compressed Sparse Banks (CSB), that transparently exposes inter-bank parallelism in BBS to hardware. Finally, we design an FPGA accelerator that takes advantage of BBS to eliminate irregular computation and memory accesses. Implemented on Intel Arria-10 FPGA, the BBS accelerator can achieve 750.9 GOPs on sparse LSTM networks with a batch size of 1. Compared to state-of-the-art FPGA accelerators for LSTM with different compression techniques, the BBS accelerator achieves 2.3 ~ 3.7x improvement on energy efficiency and 7.0 ~ 34.4x reduction on latency with negligible loss of model accuracy. Shijie Cao, Chen Zhang 0001, Zhuliang Yao, Wencong Xiao, Lanshun Nie, Dechen Zhan, Yunxin Liu 0001, Ming Wu 0007 |
FPGA | 7 |
| 2019 | DRL360: 360-degree Video Streaming with Deep Reinforcement Learningabstract360-degree videos have gained more popularity in recent years, owing to the great advance of panoramic cameras and head-mounted devices. However, as 360-degree videos are usually in high resolution, transmitting the content requires extremely high bandwidth. To protect the Quality of Experience (QoE) of users, researchers have proposed tile-based 360-degree video streaming systems that allocate high/low bit rates to selected tiles of video frames for streaming over the limited bandwidth. It is challenging to determine which tiles should be allocated with a high/low rate, because (1) the video playbacks include too many features that dynamically change over time when making the rate allocation; (2) most of the state-of-the-art systems focus on a fixed set of heuristics to optimize a specific QoE objective, while users may have various QoE objectives that need to be optimized in different ways. This paper presents a Deep Reinforcement Learning (DRL) based framework for 360-degree video streaming, named DRL360. The DRL360 framework helps improve the system performance by jointly optimizing multiple QoE objectives across a broad set of dynamic features. The DRL-based model adaptively allocates rates for the tiles of the future video frames based on the observations collected by client video players. We compare the proposed DRL360 to the existing systems by trace-driven evaluations as well as conducting a realworld experiment over a wide variety of network conditions. Evaluation results reveal that DRL360 can adapt to all considered scenarios, and outperform the state-of-the-art approaches by 20%-30% on average given different QoE objectives. Yuanxing Zhang, Kaigui Bian, Yunxin Liu 0001, Lingyang Song, Xiaoming Li 0001 |
INFOCOM | 4 |
| 2019 | Characterizing and orchestrating NFV-ready servers for efficient edge data processingabstractThe fast-growing Internet of Things (IoT) and Artificial intelligence (AI) applications mandate high-performance edge data analytics. This requirement cannot be fully fulfilled by prior works that focus on either small architectures (e.g., accelerators) or large infrastructure (e.g., cloud data centers). Sitting in between the edge and cloud, there have been many server-level designs for augmenting edge data processing. However, they often require specialized hardware resources and lack scalability as well as agility. Lu Zhang 0049, Chao Li 0009, Pengyu Wang 0003, Yunxin Liu 0001, Yang Hu 0001, Quan Chen 0002, Minyi Guo |
IWQoS | 4 |
| 2019 | HotEdgeVideo'19: Workshop on Hot Topics in Video Analytics and Intelligent EdgesabstractNo abstract available. Ganesh Ananthanarayanan, Yunxin Liu 0001, Yuanchao Shu |
MobiCom | 2 |
| 2019 | Occlumency: Privacy-preserving Remote Deep-learning Inference Using SGXabstractDeep-learning (DL) is receiving huge attention as enabling techniques for emerging mobile and IoT applications. It is a common practice to conduct DNN model-based inference using cloud services due to their high computation and memory cost. However, such a cloud-offloaded inference raises serious privacy concerns. Malicious external attackers or untrustworthy internal administrators of clouds may leak highly sensitive and private data such as image, voice and textual data. In this paper, we propose Occlumency, a novel cloud-driven solution designed to protect user privacy without compromising the benefit of using powerful cloud resources. Occlumency leverages secure SGX enclave to preserve the confidentiality and the integrity of user data throughout the entire DL inference process. DL inference in SGX enclave, however, impose a severe performance degradation due to limited physical memory space and inefficient page swapping. We designed a suite of novel techniques to accelerate DL inference inside the enclave with a limited memory size and implemented Occlumency based on Caffe. Our experiment with various DNN models shows that Occlumency improves inference speed by 3.6x compared to the baseline DL inference in SGX and achieves a secure DL inference within 72% of latency overhead compared to inference in the native environment. Taegyeong Lee, Saumay Pushp, Caihua Li, Yunxin Liu 0001, Youngki Lee 0001, Fengyuan Xu, Chenren Xu, Junehwa Song |
MobiCom | 5 |
| 2019 | A First Look at Deep Learning Apps on SmartphonesabstractTo bridge the knowledge gap between research and practice, we present the first empirical study on 16,500 the most popular Android apps, demystifying how smartphone apps exploit deep learning in the wild. To this end, we build a new static tool that dissects apps and analyzes their deep learning functions. Our study answers threefold questions: what are the early adopter apps of deep learning, what do they use deep learning for, and how do their deep learning models look like. Our study has strong implications for app developers, smartphone vendors, and deep learning R&D. On one hand, our findings paint a promising picture of deep learning for smartphones, showing the prosperity of mobile deep learning frameworks as well as the prosperity of apps building their cores atop deep learning. On the other hand, our findings urge optimizations on deep learning models deployed on smartphones, protection of these models, and validation of research ideas on these models. Mengwei Xu 0001, Yuanqiang Liu, Felix Xiaozhu Lin, Yunxin Liu 0001, Xuanzhe Liu |
WWW | 5 |
| 2018 | DeepCache: Principled Cache for Mobile Deep VisionabstractWe present DeepCache, a principled cache design for deep learning inference in continuous mobile vision. DeepCache benefits model execution efficiency by exploiting temporal locality in input video streams. It addresses a key challenge raised by mobile vision: the cache must operate under video scene variation, while trading off among cacheability, overhead, and loss in model accuracy. At the input of a model, DeepCache discovers video temporal locality by exploiting the video's internal structure, for which it borrows proven heuristics from video compression; into the model, DeepCache propagates regions of reusable results by exploiting the model's internal structure. Notably, DeepCache eschews applying video heuristics to model internals which are not pixels but high-dimensional, difficult-to-interpret data. Our implementation of DeepCache works with unmodified deep learning models, requires zero developer's manual effort, and is therefore immediately deployable on off-the-shelf mobile devices. Our experiments show that DeepCache saves inference execution time by 18% on average and up to 47%. DeepCache reduces system energy consumption by 20% on average. Mengwei Xu 0001, Mengze Zhu, Yunxin Liu 0001, Felix Xiaozhu Lin, Xuanzhe Liu |
MobiCom | 3 |
| 2018 | Cutting the Cord: Designing a High-quality Untethered VR System with Low Latency Remote RenderingabstractThis paper introduces an end-to-end untethered VR system design and open platform that can meet virtual reality latency and quality requirements at 4K resolution over a wireless link. High-quality VR systems generate graphics data at a data rate much higher than those supported by existing wireless-communication products such as Wi-Fi and 60GHz wireless communication. The necessary image encoding, makes it challenging to maintain the stringent VR latency requirements. To achieve the required latency, our system employs a Parallel Rendering and Streaming mechanism to reduce the add-on streaming latency, by pipelining the rendering, encoding, transmission and decoding procedures. Furthermore, we introduce a Remote VSync Driven Rendering technique to minimize display latency. To evaluate the system, we implement an end-to-end remote rendering platform on commodity hardware over a 60Ghz wireless network. Results show that the system can support current 2160x1200 VR resolution at 90Hz with less than 16ms end-to-end latency, and 4K resolution with 20ms latency, while keeping a visually lossless image quality to the user. Ruiguang Zhong, Wuyang Zhang, Yunxin Liu 0001, Jiansong Zhang 0001, Marco Gruteser |
MobiSys | 4 |
| 2018 | UbiTap: Leveraging Acoustic Dispersion for Ubiquitous Touch Interface on Solid SurfacesabstractWith the omnipresence of computing devices in our daily lives, interests in ubiquitous computing interfaces have grown. In response to this, various studies have introduced on-surface input techniques which use the surfaces of surrounding objects as a touch interface. However, these methods are yet struggling to support ubiquitous interaction due to their dependency on specific hardware or environments. In this paper, we propose UbiTap, an input method that turns solid surfaces into a touch input space, through the use of sound (i.e., with microphones already present in the commodity devices). More specifically, we develop a novel touch localization technique which leverages the physical phenomenon, referred to as dispersion, a characteristic of sound as it travels through solid surfaces, so as to address challenges which limit existing acoustic-based solutions in terms of portability, accuracy, usability, robustness, and responsiveness. Our extensive experiments with a prototype of UbiTap show that we can support sub-centimeter accuracy on various surfaces with minor user calibration effort. In our experience with real-world users, UbiTap significantly improves usability and robustness, thus enabling the emergence of more exciting applications. Hyosu Kim, Anish Byanjankar, Yunxin Liu 0001, Yuanchao Shu, Insik Shin |
SenSys | 3 |
| 2018 | Aladdin: Automating Release of Deep-Link APIs on AndroidabstractCompared to the Web where each web page has a global URL for external access, a specific 'page' inside a mobile app cannot be easily accessed unless the user performs several steps from the landing page of this app. Recently, the concept of 'deep link' is expected to be a promising solution and has been advocated by major service providers to enable targeting and opening a specific page of an app externally with an accessible uniform resource identifier. In this paper, we present a large-scale empirical study to investigate how deep links are really adopted, over 25,000 Android apps. To our surprise, we find that deep links have quite low coverage, e.g., more than 70% and 90% of the apps do not have deep links on app stores Wandoujia and Google Play, respectively. One underlying reason is the mandatory and non-trivial manual efforts of app developers to provide APIs for deep links. We then propose the Aladdin approach along with its supporting tool to help developers practically automate the release of deep-link APIs to access locations inside their apps. Aladdin includes a novel cooperative framework by synthesizing the static analysis and the dynamic analysis while minimally engaging developers» inputs and configurations, without requiring any coding efforts or additional deployment efforts. We evaluate Aladdin with 579 popular apps and demonstrate its effectiveness and performance. Yun Ma 0002, Ziniu Hu, Yunxin Liu 0001, Tao Xie 0001, Xuanzhe Liu |
WWW | 3 |
| 2018 | A Tale of Two Fashions: An Empirical Study on the Performance of Native Apps and Web Apps on Androidabstractprevalent smartphones have become the major entrance to accessing services on the Internet. On smartphones, users can have two options as the clients, i.e., native apps and Web apps. There have been several debates about native apps and Web apps. However, major service providers such as Google, Amazon, and Facebook provide both native apps and Web apps to end-users. Essentially, the performance differences between these two types of apps haven't been addressed. Indeed, the performance differences make non-trivial impacts on apps development, deployment, and distribution. In this article, we conduct a measurement study on the performance of native apps and Web apps on Android smartphones. Specifically, we want to explore given the same functionalities, do Web apps always perform poorly compared to native apps. We select 328 services from some popular providers, covering various domains such as e-commerce, map, social networking, and entertainment. With HTTP-level trace analysis, we demystify the workflows on how native apps and Web apps deliver services on mobile devices, respectively. Then, we characterize the performance differences between native apps and Web apps with the metrics including the number of requests, response time, data drain, and energy consumption. We find that the performance of Web apps is better than native apps in more than 31 percent cases. Our derived knowledge can suggest some recommendations to improve the performance for mobile apps. Yun Ma 0002, Xuanzhe Liu, Yi Liu 0014, Yunxin Liu 0001, Gang Huang 0001 |
IEEE Trans. Mob. Comput. | 4 |
| 2018 | Characterizing Privacy Risks of Mobile Apps with Sensitivity AnalysisabstractGiven the emerging concerns over app privacy-related risks, major app distribution providers (e.g., Microsoft) have been exploring approaches to help end users to make informed decision before installation. This is different from existing approaches of simply trusting users to make the right decision. We build on the direction of risk rating as the way to communicate app-specific privacy risks to end users. To this end, we propose to use sensitivity analysis to infer whether an app requests sensitive on-device resources/ data that are not required for its expected functionality. Our system, Privet, addresses challenges in efficiently achieving test coverage and automated privacy risk assessment. Finally, we evaluate Privet with 1,000 Android apps released in the wild. Li Lyna Zhang, Chieh-Jan Mike Liang, Zhao Lucis Li, Yunxin Liu 0001, Feng Zhao 0001, Enhong Chen |
IEEE Trans. Mob. Comput. | 4 |
| 2018 | i-Jacob: An Internetware-Oriented Approach to Optimizing Computation-Intensive Mobile Web BrowsingabstractWeb browsing is always a key requirement of Internet users. Current mobile Web apps can contain computation-intensive JavaScript logics and thus affect browsing performance. Learning from our over-decade research and development experiences of the Internetware paradigm, we present the novel and generic i - Jacob approach to improving the performance of mobile Web browsing with effective JavaScript-code offloading. Our approach proposes a programming abstraction to make mobile Web situational and adaptive to contexts, by specifying the computation-intensive and “ offloadable ” code, and develops a platform-independent lightweight runtime spanning the mobile devices and the cloud. We demonstrate the efficiency of i - Jacob with some typical computation-intensive tasks over various combinations of hardware, operating systems, browsers, and network connections. The improvements can reach up to 49× speed-up in response time and 90% saving in energy. Xuanzhe Liu, Meihua Yu, Yun Ma 0002, Gang Huang 0001, Hong Mei 0001, Yunxin Liu 0001 |
ACM Trans. Internet Techn. | 6 |
| 2018 | Learning-Aided Stochastic Network Optimization With State Prediction
Longbo Huang, Minghua Chen 0001, Yunxin Liu 0001 |
IEEE/ACM Trans. Netw. | 3 |
| 2017 | BikeLoc: a Real-time High-Precision Bicycle Localization System Using Synthetic Aperture RadarabstractIn recent years we have witnessed the rapid development of smart bicycles. For example, Mobike1 is able to interact with smartphones. As we all known, accurate bicycle localization system is one of the most critical technologies for the development of smart bicycles. However, GPS's error is at meter-level and it performs poorly under skyscrapers and in tunnels. Hongjiang Lyu, Linghe Kong, Chengzhang Li, Yunxin Liu 0001, Jiansong Zhang 0001, Guihai Chen |
APNet | 4 |
| 2017 | Latency-based WiFi congestion control in the air for dense WiFi networksabstractWiFi has become the primary method to access the Internet. However, the WiFi-hop latency, particularly in dense-WiFi environments, is far from satisfactory [1], to support delay-sensitive applications such as Web browsing and VoIP. The WiFi latency mainly comes from two kinds of queues: the host queue and the distributed queue, which is caused by CSMA/CA mechanism when multiple nodes contend for the channel. While the host queue can be easily bypassed using priority scheduling at end-host, the distributed queue is not. Previously, IEEE 802.11e tries to provide priorities in this distributed queue by adjusting the MAC layer parameters, but it does not scale when there are increasing number of delay-sensitive flows. In this paper, we propose and design QAir, a practical solution to reduce WiFi latency of delay-sensitive flows in dense WiFi networks. QAir takes a different approach to transfer this distributed queue to host queue. Consequently, the delay-sensitive flows can bypass the entire queue and their latency can be greatly reduced. QAir works in a distributed manner with no centralized scheduler. We have implemented QAir on commodity WiFi devices. Experimental results show that, compared to the 802.11 DCF baseline, QAir can reduce the average WiFi-hop latency of delay-sensitive flows by 50-75%. Changhua Pei, Youjian Zhao, Yunxin Liu 0001, Kun Tan 0002, Jiansong Zhang 0001, Yuan Meng 0002, Dan Pei |
IWQoS | 3 |
| 2017 | Enabling accurate and efficient modeling-based CPU power estimation for smartphonesabstractCPU is one of the most significant sources of power consumption on smartphones. Power modeling is a key technique and important tool for power estimation and management, both of which are critical for providing good QoS for smartphones. However, we find that existing CPU power models for smartphones are ill-suited for modern multicore CPUs: they can give high estimation errors (up to 34%) and high estimation accuracy variation (more than 30%) for different types of workloads on mainstream multicore smartphones. The cause is that the existing approaches do not appropriately consider the effects of CPU idle power states on smartphones CPU power modeling. Based on our extensive measurement experiments, we develop a new CPU power modeling approach that carefully considers the effects of CPU idle power states. We present the detailed design of our power modeling approach, and a prototype CPU power estimation system on commercial multicore smartphones. Evaluation results show that our approach consistently achieves higher power estimation accuracy and stability for various benchmarks programs and real apps than the existing approaches. Yifan Zhang 0002, Yunxin Liu 0001, Xuanzhe Liu, Qun Li 0001 |
IWQoS | 2 |
| 2017 | Systematically testing background services of mobile appsabstractContrary to popular belief, mobile apps can spend a large fraction of time running "hidden" as background services. And, bugs in services can translate into crashes, energy depletion, device slow-down, etc. Unfortunately, without necessary testing tools, developers can only resort to telemetries from user devices in the wild. To this end, Snowdrop is a testing framework that systematically identifies and automates background services in Android apps. Snowdrop realizes a service-oriented approach that does not assume all inter-component communication messages are explicitly coded in the app bytecode. Furthermore, to improve the completeness of test inputs generated, Snowdrop infers field values by exploiting the similarity in how developers name variables. We evaluate Snowdrop by testing 848 commercially available mobile apps. Empirical results show that Snowdrop can achieve 20.91% more code path coverage than pathwise test input generators, and 64.11% more coverage than random test input generators. Li Lyna Zhang, Chieh-Jan Mike Liang, Yunxin Liu 0001, Enhong Chen |
ASE | 3 |
| 2017 | RAVEN: Perception-aware Optimization of Power Consumption for Mobile GamesabstractHigh-end mobile GPUs are now becoming an integral part of mobile devices. However, a mobile GPU constitutes a major portion of power consumption on the devices, and mobile games top as the most popular class of graphics applications. This paper presents the design and implementation of RAVEN, a novel, on-the-fly frame rate scaling system for mobile gaming applications. RAVEN utilizes human visual perception of graphics change to opportunistically achieve power saving without degrading user experiences. The system develops a light-weight frame comparison technique to measure and predict perception-aware frame similarity. It also builds a low resolution virtual display which clones the device screen for performing similarity measurement at a low-power cost. It is able to work on an existing commercial smartphone and support applications from app stores without any modifications. It has been implemented on Nexus 5X, and its performance has been measured with 13 games. The system effectively reduces the overall power consumption of mobile devices while maintaining satisfactory user experiences. The power consumption is reduced by 21.78% on aver-age and up to 34.74%. Chanyou Hwang, Saumay Pushp, Changyoung Koh, Jungpil Yoon, Yunxin Liu 0001, Seungpyo Choi, Junehwa Song |
MobiCom | 5 |
| 2017 | Demo: FROG: Optimizing Power Consumption of Mobile Games Using Perception-Aware Frame Rate ScalingabstractA mobile GPU constitutes the majority of power consumption on a mobile device and mobile games top as the most popular class of graphics applications. In this demo, we present FROG, a novel frame rate optimization system for mobile gaming applications. FROG makes use of human visual perception to graphics and regulates application's frame rendering process on-the-fly for maximizing power saving without degrading the user experience. The system works on an existing commercial smartphone and support the legacy gaming applications from app stores without requiring any changes from applications Saumay Pushp, Chanyou Hwang, Changyoung Koh, Jungpil Yoon, Yunxin Liu 0001, Seungpyo Choi, Junehwa Song |
MobiCom | 5 |
| 2017 | Learning-aided Stochastic Network Optimization with Imperfect State PredictionabstractWe investigate the problem of stochastic network optimization in the presence of imperfect state prediction and non-stationarity. Based on a novel distribution-accuracy curve prediction model, we develop the predictive learning-aided control (PLC) algorithm, which jointly utilizes historic and predicted network state information for decision making. PLC is an online algorithm that requires zero a-prior system statistical information, and consists of three key components, namely sequential distribution estimation and change detection, dual learning, and online queue-based control. Longbo Huang, Minghua Chen 0001, Yunxin Liu 0001 |
MobiHoc | 3 |
| 2017 | BikeMate: Bike Riding Behavior Monitoring with SmartphonesabstractDetecting dangerous riding behaviors is of great importance to improve bicycling safety. Existing bike safety precautionary measures rely on dedicated infrastructures that incur high installation costs. In this work, we propose BikeMate, a ubiquitous bicycling behavior monitoring system with smartphones. BikeMate invokes smartphone sensors to infer dangerous riding behaviors including lane weaving, standing pedalling and wrong-way riding. For easy adoption, BikeMate leverages transfer learning to reduce the overhead of training models for different users, and applies crowdsourcing to infer legal riding directions without prior knowledge. Experiments with 12 participants show that BikeMate achieves an overall accuracy of 86.8% for lane weaving and standing pedalling detection, and yields a detection accuracy of 90% for wrong-way riding using crowdsourced GPS traces. Weixi Gu, Zimu Zhou, Yuxun Zhou, Han Zou, Yunxin Liu 0001, Costas J. Spanos, Lin Zhang 0001 |
MobiQuitous | 5 |
| 2017 | AppHolmes: Detecting and Characterizing App Collusion among Third-Party Android MarketsabstractBackground activities on smartphones are essential to today's "always-on" mobile device experience. Yet, there lacks a clear understanding of the cooperative behaviors among background activities as well as a quantification of the consequences. In this paper, we present the first in-depth study of app collusion, in which one app surreptitiously launches others in the background without user's awareness. To enable the study, we develop AppHolmes, a static analysis tool for detecting app collusion by examining the app binaries. By analyzing 10,000 apps from top third-party app markets, we found that i) covert, cooperative behaviors in background app launch are surprisingly pervasive, ii) most collusion is caused by shared services, libraries, or common interest among apps, and iii) collusion has serious impact on performance, efficiency, and security. Overall, our work presents a strong implication on future mobile system design. Mengwei Xu 0001, Yun Ma 0002, Xuanzhe Liu, Felix Xiaozhu Lin, Yunxin Liu 0001 |
WWW | 5 |
| 2017 | ShuffleDog: Characterizing and Adapting User-Perceived Latency of Android AppsabstractNumerous complains have been made by Android users who severely suffer from the sluggish response when interacting with their devices. However, very few studies have been conducted to understand the user-perceived latency or mitigate the UI-lagging problem. In this paper, we conduct the first systematic measurement study to quantify the user-perceived latency using typical interaction-intensive Android apps in running with and without background workloads. We reveal the insufficiency of Android system in ensuring the performance of foreground apps and therefore design a new system to address the insufficiency accordingly. We develop a lightweight tracker to accurately identify all delay-critical threads that contribute to the slow response of user interactions. We then build a resource manager that can efficiently schedule various system resources including CPU, I/O, and GPU, for optimizing the performance of these threads. We implement the proposed system on commercial smartphones and conduct comprehensive experiments to evaluate our implementation. Evaluation results show that our system is able to significantly reduce the user-perceived latency of foreground apps in running with aggressive background workloads, up to 10x, while incurring negligible system overhead of less than 3.1 percent CPU and 7 MB memory. Gang Huang 0001, Mengwei Xu 0001, Felix Xiaozhu Lin, Yunxin Liu 0001, Yun Ma 0002, Saumay Pushp, Xuanzhe Liu |
IEEE Trans. Mob. Comput. | 4 |
| 2017 | ReWAP: Reducing Redundant Transfers for Mobile Web Browsing via App-Specific Resource PackagingabstractRedundant transfer of resources is a critical issue for compromising the performance of mobile Web applications (a.k.a., apps) in terms of data traffic, load time, and even energy consumption. Evidence demonstrates that the current cache mechanisms are far from satisfactory. With lessons learned from how native apps manage their resources, in this article, we present the ReWAP approach to fundamentally reducing redundant transfers by restructuring the resource loading of mobile Web apps. ReWAP is based on an efficient resource-packaging mechanism where stable resources are encapsulated and maintained into a package, and such a package shall be loaded always from the local storage and updated by explicitly refreshing. By retrieving and analyzing the update of resources, ReWAP maintains resource packages that can accurately identify which resources can be loaded from the local storage for a considerably long period. ReWAP also provides a wrapper for mobile Web apps to enable loading and updating resource packages in the local storage as well as loading resources from resource packages. ReWAP can be easily and seamlessly deployed into existing mobile Web architectures with minimal modifications, and is transparent to end-users. We evaluate ReWAP based on continuous 15day access traces of 50 mobile Web apps randomly chosen from Alexa top 500 ranking list. Compared to the original mobile Web apps with cache enabled, ReWAP can significantly reduce the data traffic, with the median saving up to 51 percent. In addition, ReWAP can incur only very minor runtime overhead of the client-side browsers and thus does not compromise user experiences. Xuanzhe Liu, Yun Ma 0002, Shuailiang Dong, Yunxin Liu 0001, Tao Xie 0001, Gang Huang 0001 |
IEEE Trans. Mob. Comput. | 4 |
| 2017 | SWAROVsky: Optimizing Resource Loading for Mobile Web BrowsingabstractImperfect Web resource loading prevents mobile Web browsing from providing satisfactory user experience. In this article, we design and implement the SWAROVsky system to address three main issues of current inefficient Web resource loading: (1) on-demand and thus slow loading of sub-resources of webpages; (2) duplicated loading of resources with different URLs but the same content; and (3) redundant loading of the same resource due to improper cache configurations. SWAROVsky employs a dual-proxy architecture that comprises a remote cloud-side proxy and a local proxy on mobile devices. The remote proxy proactively loads webpages from their original Web servers and maintains a resource loading graph for every single webpage. Based on the graph, the remote proxy is capable of deciding which resources are “really” needed for the webpage and their loading orders, and thus can synchronize these needed resources with the local proxy of a client efficiently and timely. The local proxy also runs an intelligent and light-weight algorithm to identify resources with different URLs but the same content, and thus can avoid duplicated downloading of the same content via network. Our system can be used with existing Web browsers and Web servers, and does not break the normal semantics of a webpage. Evaluations with 50 websites show that on average our system can reduce the page load time by 43.1 percent and the network data transmission by 57.6 percent, while imposing marginal system overhead. Xuanzhe Liu, Yun Ma 0002, Yunxin Liu 0001, Tao Xie 0001, Gang Huang 0001 |
IEEE Trans. Mob. Comput. | 4 |
| 2016 | ARTcode: preserve art and code in any imageabstractThe ubiquitous QR codes and some similar barcodes are becoming a convenient and popular approach to impromptu communication between mobile devices and their surrounding cyber-physical world. However, such codes suffer from two common drawbacks: poor viewing experience and inability to be identified through itself. In this work, we propose ART-code-- Adaptive Robust doT matrix barcode, which aims to preserve ART and CODE features in one visual pattern. It works on any surface (paper or electronic displays) and is able to convert any image or any form of human-readable contents (e.g., a picture, a logo, a slogan) into an ARTcode. It looks like an image which retains human-readable and aesthetically pleasant contents, and in the meanwhile, it acts as a QR code which conveys data bits over the visual channel. The core enablers in ARTcode are (1) the design of the colored dot matrix for data embedding with little distortion from the original image and (2) a comprehensive error correction scheme which enhances decoding robustness against noises and interferences from the original image in ARTcode. We implement ARTcode with the receiver on Android phones and the sender from a PC or a phone (it can be printed in paper). We conduct extensive user survey and experiments for evaluation. It validates the effectiveness and wide applicability of ARTcode: It works well with all of 197 images randomly downloaded, covering representative categories of the gray-scale images, logos, colored ones with low/medium/strong contrasts. The image quality is quite acceptable in a subjective user-perception survey with 50 participants and data communication accuracy achieves as high as 99% in almost all the cases (> 96% raw accuracy in ARTcode without error detection and other schemes). Yuting Bao, Chuhao Luo, Xingya Zhao, Chunyi Peng 0001, Yunxin Liu 0001, Xinbing Wang |
UbiComp | 7 |
| 2016 | AMIL: Localizing neighboring mobile devices through a simple gestureabstractSmartphone users are often grouped to exchange files or perform collaborative tasks when meeting together. We argue that the location information of group members is critical to many mobile applications. Existing localization solutions mostly rely on anchor nodes or infrastructures to perform ranging and positioning. These approaches are inefficient for ad hoc scenarios. In this paper, we propose AMIL, an Acoustic Mobility-Induced TDoA (Time-Difference-of-Arrival)-based Localization scheme for smartphones. In AMIL, a smartphone user can use simple gestures (e.g., hold the phone and draw a triangle in the air) to quickly obtain the relative coordinates of neighboring mobile devices. We have implemented and evaluated AMIL on off-the-shelf smartphones. The field tests have shown that our scheme can achieve less than three degree orientation errors and can successfully build a simple map of 12 people in an office room with average error of 50cm. Shanhe Yi, Qun Li 0001, Guobin Shen, Yunxin Liu 0001, Edmund Novak |
INFOCOM | 5 |
| 2016 | Demystifying the Imperfect Client-Side Cache Performance of Mobile Web BrowsingabstractThe web browser is one of the most significant applications on mobile devices such as smartphones. However, the user experience of mobile web browsing is undesirable because of the slow resource loading. To improve the performance of web resource loading, client-side cache has been adopted as a key mechanism. However, the existing passive measurement studies cannot comprehensively characterize the “client-side” cache performance of mobile web browsing. For example, most of these studies mainly focus on client-side implementations but not server-side configurations, suffer from biased user behaviors, and fail to study “miscached” resources. To address these issues, in this article, we present a proactive approach to making a comprehensive measurement study on client-side cache performance. The key idea of our approach is to proactively crawl resources from hundreds of websites periodically with a fine-grained time interval. Thus, we are able to uncover the resource update history and cache configurations at the server side, and analyze the cache performance in various time granularities. Based on our collected data, we build a new cache analysis model and study the upper bound of how high percentage of resources could potentially be cached and how effectively the caching works in practice. We report detailed analysis results of different websites and various types of web resources, and identify the problems caused by unsatisfactory cache performance. In particular, we identify two major problems - Redundant Transfer and Miscached Resource, which lead to unsatisfactory cache performance. We investigate three main root causes: Same Content, Heuristic Expiration, and Conservative Expiration Time, and discuss what mobile web developers can do to mitigate those problems. Xuanzhe Liu, Yun Ma 0002, Yunxin Liu 0001, Tao Xie 0001, Gang Huang 0001 |
IEEE Trans. Mob. Comput. | 3 |
| 2015 | Characterizing RESTful Web Services Usage on Smartphones: A Tale of Native Apps and Web AppsabstractThe burst of Web-based Restful services brings us a number of facilities in our life and work. We are used to take smartphones to access these Web services, like location-based services, weather search, mapping, social networking, et al. On smartphones, we have two options of service consumers, a.k.a, Native apps and Web apps. Despite the platform-independence, Web apps are claimed to provide the same features and comparable user experiences with native apps. However, one fact is that more and more people prefer native apps rather than Web apps. In this paper, we make an empirical study on characterizing the performance disparity of native apps and Web apps. Given the same functionalities provided by the same service providers, we explore the Restful Web services that are used by native apps and Web apps. With HTTP-level trace analysis, we demystify the workflows on how native apps and Web apps use Web services and summarize different service usage patterns from architectural style perspective. Then we characterize the performance differences between native apps and Web apps on realizing Restful Web services including GET, DELETE, PUT & POST, in terms of number of network connections, response time, and data drain, given the same functional features. Our observations reveal that Web apps do not always perform worse than native apps using Restful Web services under the same context. We further propose some implications to improve both native apps and Web apps on smartphones. Yi Liu 0014, Xuanzhe Liu, Yun Ma 0002, Yunxin Liu 0001, Zibin Zheng, Gang Huang 0001, M. Brian Blake |
ICWS | 4 |
| 2015 | Mash Droid: An Approach to Mobile-Oriented Dynamic Services Discovery and Composition by In-App SearchabstractThe popularity of smartphones and tablet computers in recent years makes mobile apps burst. Mobile apps have become the main consumers of the Internet-based services. Compared to traditional applications in the desktop computing era, mobile devices with their apps bring new opportunities and challenges to service computing community, in various aspects like service publication, discovery, interaction, composition, et al. In this paper, we propose a novel data-driven, content-based mobile apps composition approach, called Mash Droid, by leveraging a novel In-App Search mechanism, i.e., Discovering relevant services for the data and content in apps. Rather than existing techniques that usually integrate fixed Web services, our approach relies on the dynamic service discovery and flexible data exchange between several apps. The unique feature of our approach is enabling the data communication channel between apps by the content index services provided by a leading Android appstore, Wandoujia, which now has over 1,000,000 apps and 200 million users. We employ the In-App Search mechanism to define a Restful-style app model and resource-oriented app description model. Based on the models, we design a framework for dynamically discovering relevant apps that could be composed with current app's contexts. We implement a prototype to demonstrate our approach. Yun Ma 0002, Xuanzhe Liu, Meihua Yu, Yunxin Liu 0001, Qiaozhu Mei, Feng Feng 0001 |
ICWS | 4 |
| 2015 | Rethinking Energy-Performance Trade-Off in Mobile Web Page LoadingabstractWeb browsing is a key application on mobile devices. However, mobile browsers are largely optimized for performance, imposing a significant burden on power-hungry mobile devices. In this work, we aim to reduce the energy consumed to load web pages on smartphones, preferably without increasing page load time and compromising user experience. To this end, we first study the internals of web page loading on smartphones and identify its energy-inefficient behaviors. Based on our findings, we then derive general design principles for energy-efficient web page loading, and apply these principles to the open-source Chromium browser and implement our techniques on commercial smartphones. Experimental results show that our techniques are able to achieve a 24.4% average system energy saving for Chromium on a latest-generation big.LITTLE smartphone using WiFi (a 22.5% saving when using 3G), while not increasing average page load time. We also show that our proposed techniques can bring a 10.5% system energy saving on average with a small 1.69\% increase in page load time for mobile Firefox web browser. User study results indicate that such a small increase in page load time is hardly perceivable. Duc Hoang Bui, Yunxin Liu 0001, Hyosu Kim, Insik Shin, Feng Zhao 0001 |
MobiCom | 2 |
| 2015 | Optimizing Smartphone Power Consumption through Dynamic Resolution ScalingabstractThe extremely-high display density of modern smartphones imposes a significant burden on power consumption, yet does not always provide an improved user experience and may even lead to a compromised user experience. As human visually-perceivable ability highly depends on the user-screen distance, a reduced display resolution may still achieve the same user experience when the user-screen distance is large. This provides new power-saving opportunities. In this paper, we present a flexible dynamic resolution scaling system for smartphones. The system adopts an ultrasonic-based approach to accurately detect the user-screen distance at low-power cost and makes scaling decisions automatically for maximum user experience and power saving. App developers or users can also adjust the resolution manually as their needs. Our system is able to work on existing commercial smartphones and support legacy apps, without requiring re-building the ROM or any changes of apps. An end-to-end dynamic resolution scaling system is implemented on the Galaxy S5 LTE-A and Nexus 6 smartphones, and the correctness and effectiveness are evaluated against 30 games and benchmarks. Experimental results show that all the 30 apps can run successfully with per-frame, real-time dynamic resolution scaling. The energy per frame can be reduced by 30.1% on average and up to 60.5\% at most when the resolution is halved, for 15 apps. A user study with 10 users indicates that our system remains good user experience, as none of the 10 users could perceive the resolution changes in the user study. Songtao He, Yunxin Liu 0001, Hucheng Zhou |
MobiCom | 2 |
| 2015 | Demo: Optimizing Smartphone Power Consumption through Dynamic Resolution ScalingabstractThe extremely-high display density of modern smartphones imposes a significant burden on power consumption, yet does not always provide an improved user experience and may even lead to a compromised user experience. As human visually-perceivable ability highly depends on the user-screen distance, a reduced display resolution may still achieve the same user experience when the user-screen distance is large. This provides new power-saving opportunities. We present a flexible dynamic resolution scaling system for smartphones. The system adopts an ultrasonic-based approach to detect the user-screen distance at low-power cost and makes scaling decisions automatically for maximum user experience and power saving. App developers or users can also adjust the resolution manually and dynamically as their needs. Our system is able to work on the existing commercial smartphones and support the legacy apps, without requiring re-building the ROM or any changes from apps. Songtao He, Yunxin Liu 0001, Hucheng Zhou |
MobiCom | 2 |
| 2015 | EarlyBird: Mobile Prefetching of Social Network Feeds via Content Preference Mining and Usage Pattern AnalysisabstractSocial networks are the most engaging applications on mobile devices, and they are becoming the main sources for users to consume content. However, content retrieval, especially for embedded links and multimedia, can often be too slow, too energy hungry or too expensive for on-the-go mobile users. To address these issues, we collect and analyze a large set of traces from over 6000 real-life users of a popular mobile Twitter client. Based on the unique challenges identified from our dataset, we present inference-based social network content prefetcher, Earlybird. It uses the specific signals unique to social data in order to retrieve news feeds and associated links and multimedia ahead of users' usage. Our regression-based content prediction model is able to estimate a user's likely content interests 55% of the time. Second, we develop a prefetch scheduling scheme to maximize delay reduction under users' resource constraints. For validation, we apply Earlybird to our collected dataset. We show that on average users can reduce their delays by 62% at the cost of no more than 3% battery and 40MB/month cellular data. Xin Liu 0002, David Chu, Yunxin Liu 0001 |
MobiHoc | 4 |
| 2015 | Measurement and Analysis of Mobile Web Cache PerformanceabstractThe Web browser is a killer app on mobile devices such as smartphones. However, the user experience of mobile Web browsing is undesirable because of the slow resource loading. To improve the performance of Web resource loading, caching has been adopted as a key mechanism. However, the existing passive measurement studies cannot comprehensively characterize the performance of mobile Web caching. For example, most of these studies mainly focus on client-side implementations but not server-side configurations, suffer from biased user behaviors, and fail to study "miscached" resources. To address these issues, in this paper, we present a proactive approach for a comprehensive measurement study on mobile Web cache performance. The key idea of our approach is to proactively crawl resources from hundreds of websites periodically with a fine-grained time interval. Thus, we are able to uncover the resource update history and cache configurations at the server side, and analyze the cache performance in various time granularities. Based on our collected data, we build a new cache analysis model and study the upper bound of how high percentage of resources could potentially be cached and how effective the caching works in practice. We report detailed analysis results of different websites and various types of Web resources, and identify the problems caused by unsatisfactory cache performance. In particular, we identify two major problems -- Redundant Transfer and Miscached Resource, which lead to unsatisfactory cache performance. We investigate three main root causes: Same Content, Heuristic Expiration, and Conservative Expiration Time, and discuss what mobile Web developers can do to mitigate those problems. Yun Ma 0002, Xuanzhe Liu, Ruirui Xiang, Yunxin Liu 0001, Tao Xie 0001 |
WWW | 5 |
| 2015 | Data-Driven Composition for Service-Oriented Situational Web ApplicationsabstractThe convergence of Services Computing and Web 2.0 gains a large space of opportunities to compose “situational” web applications from web-delivered services. However, the large number of services and the complexity of composition constraints make manual composition difficult to application developers, who might be non-professional programmers or even end-users. This paper presents a systematic data-driven approach to assisting situational application development. We first propose a technique to extract useful information from multiple sources to abstract service capabilities with a set tags. This supports intuitive expression of user's desired composition goals by simple queries, without having to know underlying technical details. A planning technique then exploits composition solutions which can constitute the desired goals, even with some potential new interesting composition opportunities. A browser-based tool facilitates visual and iterative refinement of composition solutions, to finally come up with the satisfying outputs. A series of experiments demonstrate the efficiency and effectiveness of our approach. Xuanzhe Liu, Yun Ma 0002, Gang Huang 0001, Junfeng Zhao 0001, Hong Mei 0001, Yunxin Liu 0001 |
IEEE Trans. Serv. Comput. | 6 |
| 2014 | Hidden costs of mobile data access in cellular networks (Invited paper)abstractIn mobile data access, cellular operators usually charge users purely based on the delivered data volume. However, the actual used radio resources in transmitting the same amount of data might be significantly different in various traffic patterns. In this paper, we conduct a in-depth study to unveil the radio resource consumption for mobile data transfer. We disclose that radio resources are consumed not only for data transfer but also for control signals and the overhead of channel allocation. We further develop an approach to quantify the total radio resource costs, particularly the hidden costs of signaling messages and dedicated channels. We have conducted measurement of various mobile apps in two operational cellular networks. The results reveal huge hidden costs in some always-online apps and demonstrate that apps diversity yields unproportional radio resource usage with a gap as large as 116 times. Chunyi Peng 0001, Yunxin Liu 0001 |
LANMAN | 3 |
| 2014 | Design, Realization, and Evaluation of DozyAP for Power-Efficient Wi-Fi TetheringabstractWi-Fi tethering (i.e., sharing the Internet connection of a mobile phone via its Wi-Fi interface) is a useful functionality and is widely supported on commercial smartphones. Yet, existing Wi-Fi tethering schemes consume excessive power: they keep the Wi-Fi interface in a high power state regardless if there is ongoing traffic or not. In this paper, we propose DozyAP to improve the power efficiency of Wi-Fi tethering. Based on measurements in typical applications, we identify many opportunities that a tethering phone could sleep to save power. We design a simple yet reliable sleep protocol to coordinate the sleep schedule of the tethering phone with its clients without requiring tight time synchronization. Furthermore, we develop a two-stage, sleep interval adaptation algorithm to automatically adapt the sleep intervals to ongoing traffic patterns of various applications. DozyAP does not require any changes to the 802.11 protocol and is incrementally deployable through software updates. We have implemented DozyAP on commercial smartphones. Experimental results show that, while retaining comparable user experiences, our implementation can allow the Wi-Fi interface to sleep for up to 88% of the total time in several different applications and reduce the system power consumption by up to 33% under the restricted programmability of current Wi-Fi hardware. Yunxin Liu 0001, Guobin Shen, Yongguang Zhang, Qun Li 0001, Chiu C. Tan 0001 |
IEEE/ACM Trans. Netw. | 2 |
| 2013 | Content-based isolation: rethinking isolation policy design on client systemsabstractModern client platforms, such as iOS, Android, Windows Phone, and Windows 8, have progressed from a per-user isolation policy, where users are isolated but a user's applications run in the same isolation container, to an application isolation policy, where different applications are isolated from one another. However, this is not enough because mutually distrusting content can interfere with one another inside a single application. For example, an attacker-crafted image may compromise a photo editor application and steal other images processed by the editor. Alexander Moshchuk, Helen J. Wang, Yunxin Liu 0001 |
CCS | 3 |
| 2013 | AppMobiCloud: improving mobile web applications by mobile-cloud convergenceabstractBenefitting from advanced web technologies like JavaScript, CSS3 and HTML5, current web applications can provide ever richer functionalities and user experiences, on both PC and mobile devices like tablet computers and smartphones. Furthermore, they can perform complex computations which are usually resource-intensive and consuming, e.g., data analytic application and augmented reality games. Mobile devices might suffer from their limited computing capabilities and resources. As mobile devices now are gaining access through excellent connectivity with much more powerful cloud-side services, and offloading can be a potential solution. This paper presents the design and implementation of the AppMobiCloud system for improving mobile web applications by leveraging the mobile-cloud convergence. At development time, AppMobiCloud employs a combination of profiling and points-to analysis. This facilitates application developers to find the computation-intensive code fragments, and specifies whether they can be offloaded with some constraints. At runtime, AppMobiCloud migrates the chosen JavaScript code fragments from the mobile devices for remote execution. It synchronizes client-side application runtime context and constructs the "cloned" context at server, executing the codes there and re-integrating the result back to the mobile device. We evaluate our approach on three well-known JavaScript benchmarks, Dromaeo, V8 and Kraken, and a typical computation-intensive AI game. The evaluation demonstrates that our work can reduce JavaScript application's execution time and energy consumption respectively on mobile devices up to 98% and 83%. Xuanzhe Liu, Gang Huang 0001, Yunxin Liu 0001 |
Internetware | 4 |
| 2013 | MoodScope: building a mood sensor from smartphone usage patternsabstractWe report a first-of-its-kind smartphone software system, MoodScope, which infers the mood of its user based on how the smartphone is used. Compared to smartphone sensors that measure acceleration, light, and other physical properties, MoodScope is a "sensor" that measures the mental state of the user and provides mood as an important input to context-aware computing. We run a formative statistical mood study with smartphone-logged data collected from 32 participants over two months. Through the study, we find that by analyzing communication history and application usage patterns, we can statistically infer a user's daily mood average with an initial accuracy of 66%, which gradu-ally improves to an accuracy of 93% after a two-month personal-ized training period. Motivated by these results, we build a service, MoodScope, which analyzes usage history to act as a sensor of the user's mood. We provide a MoodScope API for developers to use our system to create mood-enabled applications. We further create and deploy a mood-sharing social application. Robert LiKamWa, Yunxin Liu 0001, Nicholas D. Lane, Lin Zhong 0001 |
MobiSys | 2 |
| 2013 | MoodScope: building a mood sensor from smartphone usage patternsabstractWe present MoodScope, a software system which infers the mood of its user based on how the smartphone is used. Similar to smartphone sensors that measure acceleration, light, and other physical properties, MoodScope is a "sensor" that measures the mental state of the user and provides mood as an important input to context-aware computing. We run a formative statistical study with smartphone-logged data collected from 32 participants over two months. Through the study, we find that by analyzing communication history and application usage patterns, we can statistically infer a user's daily mood average with an accuracy of 93% after a two-month training period. Motivated by these results, we build a service, MoodScope, which analyzes usage history to act as a sensor of the user's mood. Robert LiKamWa, Yunxin Liu 0001, Nicholas D. Lane, Lin Zhong 0001 |
MobiSys | 2 |
| 2013 | Optimizing background email sync on smartphonesabstractEmail is a key application used on smartphones. Even when the phone is in stand-by mode, users expect the phone to continue syncing with an email server to receive new mes-sages. Each such sync operation wakes up the smartphone for data reception and processing. In this paper, we show that this "cost of email sync" in stand-by mode constitutes a significant source of energy consumption, and thus reduces battery life. We quantify the power performance of different existing email clients on two smartphone platforms, An-droid and Windows Phone, and study the impact of system parameters such as email size, inbox size, and pull vs. push. Our results show that existing email clients do not handle email sync in an energy efficient way. This is because the underlying protocols and architectures are not designed for the specific needs of operating in stand-by mode. Based on our findings, we derive general design principles for energy-efficient event handling on smartphones, and apply these principles to the case of email sync and implement our techniques on commercial smartphones. Experimental results show that our techniques are able to significantly reduce energy cost of email sync by 49.9% on average with our experiment settings. Fengyuan Xu, Yunxin Liu 0001, Thomas Moscibroda, Ranveer Chandra, Yongguang Zhang, Qun Li 0001 |
MobiSys | 2 |
| 2013 | V-edge: Fast Self-constructive Power Modeling of Smartphones Based on Battery Voltage Dynamics
Fengyuan Xu, Yunxin Liu 0001, Qun Li 0001, Yongguang Zhang |
NSDI | 2 |
| 2013 | IP-Geolocation Mapping for Moderately Connected Internet RegionsabstractMost IP-geolocation mapping schemes [14], [16], [17], [18] take delay-measurement approach, based on the assumption of a strong correlation between networking delay and geographical distance between the targeted client and the landmarks. In this paper, however, we investigate a large region of moderately connected Internet and find the delay-distance correlation is weak. But we discover a more probable rule - with high probability the shortest delay comes from the closest distance. Based on this closest-shortest rule, we develop a simple and novel IP-geolocation mapping scheme for moderately connected Internet regions, called GeoGet. In GeoGet, we take a large number of webservers as passive landmarks and map a targeted client to the geolocation of the landmark that has the shortest delay. We further use JavaScript at targeted clients to generate HTTP/Get probing for delay measurement. To control the measurement cost, we adopt a multistep probing method to refine the geolocation of a targeted client, finally to city level. The evaluation results show that when probing about 100 landmarks, GeoGet correctly maps 35.4 percent clients to city level, which outperforms current schemes such as GeoLim [16] and GeoPing [14] by 270 and 239 percent, respectively, and the median error distance in GeoGet is around 120 km, outperforming GeoLim and GeoPing by 37 and 70 percent, respectively. Dan Li 0001, Chuanxiong Guo, Yunxin Liu 0001, Zhi-Li Zhang, Yongguang Zhang |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2012 | DozyAP: power-efficient Wi-Fi tetheringabstractWi-Fi tethering (i.e., sharing the Internet connection of a mobile phone via its Wi-Fi interface) is a useful functionality and is widely supported on commercial smartphones. Yet existing Wi-Fi tethering schemes consume excessive power: they keep the Wi-Fi interface in a high power state regardless if there is ongoing traffic or not. In this paper we propose DozyAP to improve the power efficiency of Wi-Fi tethering. Based on measurements in typical applications, we identify many opportunities that a tethering phone could sleep to save power. We design a simple yet reliable sleep protocol to coordinate the sleep schedule of the tethering phone with its clients without requiring tight time synchronization. Furthermore, we develop a two-stage, sleep interval adaptation algorithm to automatically adapt the sleep intervals to ongoing traffic patterns of various applications. DozyAP does not require any changes to the 802.11 protocol and is incrementally deployable through software updates. We have implemented DozyAP on commercial smartphones. Experimental results show that, while retaining comparable user experiences, our implementation can allow the Wi-Fi interface to sleep for up to 88% of the total time in several different applications, and reduce the system power consumption by up to 33% under the restricted programmability of current Wi-Fi hardware. Yunxin Liu 0001, Guobin Shen, Yongguang Zhang, Qun Li 0001 |
MobiSys | 2 |
| 2010 | Design, Realization, and Evaluation of xShare for Impromptu Sharing of Mobile PhonesabstractMobile phones are truly personal devices loaded with personal data such as photos, contacts, and call history. Yet it is often necessary or desirable to share our phones with others. This is especially true as mobile phones are integrating features conventionally provided by other dedicated devices, from MP3 players to games consoles. Yet existing phones assume a single user and provide little protection for private data and applications when a phone is shared. That is, when we lend our phones to others, we give away complete access. In this work, we present xShare, a protection solution to address this problem. xShare allows phone owners to rapidly specify what they want to share and place the phone into a restricted mode where only the data and applications intended for sharing can be accessed. We first present two formative user studies and derive the design requirements of xShare. We then offer the design of xShare based on file-level access control. We describe the implementation of xShare on Windows Mobile and report a comprehensive evaluation, including performance measurements, usability, and a one-month field trial. Yunxin Liu 0001, Ahmad Rahmati, Hyukjae Jang, Yuanhe Huang, Lin Zhong 0001, Yongguang Zhang, Shensheng Zhang |
IEEE Trans. Mob. Comput. | 1 |
| 2009 | Mining the Web and the Internet for Accurate IP Address GeolocationsabstractIn this paper, we present Structon, a novel approach that uses Web mining together with inference and IP traceroute to geolocate IP addresses with significantly better accuracy than existing automated approaches. Structon is composed of three ideas which we realize in three corresponding steps. First, we extract geolocation information of Web server IP addresses from Web pages. Second, we devise heuristic algorithms to improve both the accuracy and the coverage of the IP geolocation database using these Web server IP addresses and their geolocations as input. Third, for those segments that are not covered in the first two steps, we use IP traceroute to identify the access routers of those segments. When the location of the access router is known, we can deduce the location of the associated segment since it is co-located together with the access router. By mining 500-million Web pages collected in China in 2006 (11 percent of the total Web pages in China at that time), we are able to identify the geolocations for 103 million IP addresses. This represents nearly 88 percent IP addresses allocated to China in March 2008. Structon is 87.4 percent accurate at city granularity and up to 93.5 percent accurate at province level. We also used 10 day Windows Live client log to evaluate our client IP addresses coverage: Structon identified geolocations of 98.9 percent of client IP addresses. Chuanxiong Guo, Yunxin Liu 0001, Wenchao Shen, Helen J. Wang, Yongguang Zhang |
INFOCOM | 2 |
| 2009 | xShare: supporting impromptu sharing of mobile phonesabstractLoaded with personal data, e.g. photos, contacts, and call history, mobile phones are truly personal devices. Yet it is often necessary or desirable to share our phones with others. This is especially true as mobile phones are integrating features conventionally provided by other dedicated devices, from MP3 players to games consoles. Unfortunately, when we lend our phones to others, we give away complete access because existing phones assume a single user and provide little protection for private data and applications. In this work, we present xShare, a protection solution to address this problem. xShare allows phone owners to rapidly specify what they want to share and place the phone into a restricted mode where only the data and applications intended for sharing can be accessed. Yunxin Liu 0001, Ahmad Rahmati, Yuanhe Huang, Hyukjae Jang, Lin Zhong 0001, Yongguang Zhang, Shensheng Zhang |
MobiSys | 1 |