VLDB 2026 Research / reviewers in the wild / expert
Shiqi Jiang 0002
dblp:07/10820-2
· DBLP profile ↗
21ranked-venue papers
3as first author
16since 2021 · last 2026
0000-0002-4685-9633ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 16 · 3 first-author · 13 since 2021Systems, architecture and hardware · 2 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Scaling LLM Test-Time Compute with Mobile NPU on SmartphonesabstractDeploying Large Language Models (LLMs) on mobile devices faces the challenge of insufficient performance in smaller models and excessive resource consumption in larger ones. This paper highlights that mobile Neural Processing Units (NPUs) have underutilized computational resources, particularly their matrix multiplication units, during typical LLM inference. To leverage this wasted compute capacity, we propose applying parallel test-time scaling techniques on mobile NPUs to enhance the performance of smaller LLMs. However, this approach confronts inherent NPU challenges, including inadequate hardware support for fine-grained quantization and low efficiency in general-purpose computations. To overcome these, we introduce two key techniques: a hardware-aware tile quantization scheme that aligns group quantization with NPU memory access patterns, and efficient LUT-based replacements for complex operations such as Softmax and dequantization. We design and implement an end-to-end inference system that leverages the NPU's compute capability to support test-time scaling on Qualcomm Snapdragon platforms. Experiments show our approach brings significant speedups: up to 19.0× for mixed-precision GEMM and 2.2× for Softmax. More importantly, we demonstrate that smaller models using test-time scaling can match or exceed the accuracy of larger models, achieving a new performance-cost Pareto frontier. Zixu Hao, Jianyu Wei, Tuowei Wang, Minxing Huang, Huiqiang Jiang, Shiqi Jiang 0002, Ting Cao 0003, Ju Ren 0001 |
EuroSys | 6 |
| 2026 | AVA: Towards Agentic Video Analytics with Vision Language Models
Yuxuan Yan, Shiqi Jiang 0002, Ting Cao 0003, Yifan Yang 0004, Qianqian Yang 0002, Yuanchao Shu, Yuqing Yang 0001, Lili Qiu |
NSDI | 2 |
| 2025 | StreamMind: Unlocking Full Frame Rate Streaming Video Dialogue through Event-Gated Cognition
Hao Wu 0067, Yifan Yang 0004, Shiqi Jiang 0002, Qianxi Zhang, Donglin Bai, Zhibo Chen 0001, Ting Cao 0003 |
ICCV | 4 |
| 2025 | Demo: EdgeMind-OS: A Plug-and-Play Embodied Intelligence System for Real-Time On-Device DeploymentabstractBuilding an always-on, contextual AI assistant that proactively supports humans remains a central goal in Embodied AI—yet cloud-based pipelines struggle to meet due to delay, bandwidth, and privacy constraints. This demo presents EdgeMind-OS, a fully on-device intelligence system designed for embodied agents operating in real-world scenarios. Edge-Mind-OS features a hierarchical architecture combining a real-time StreamBrain, modular skill experts, and a dynamic scene-episode memory. Achieving up to 7.3× faster local processing, it enables low-latency, privacy-preserving, and plug-and-play deployment across tasks such as semantic navigation, spatial memory recall, and multimodal interaction. We demonstrate how EdgeMind-OS empowers a mobile robot with only basic locomotion capabilities to perform realtime, free-form user-robot interaction through autonomous perception, reasoning and action —without reliance on external cloud infrastructure. Jianyu Wei, Fucheng Jia, Liang Mi, Ruofei Ju, Xianye Wang, Yikai Zheng, Weijun Wang 0001, Shiqi Jiang 0002, Yunxin Liu 0001, Ting Cao 0003 |
MobiCom | 10 |
| 2025 | Babel: A Scalable Pre-trained Model for Multi-Modal Sensing via Expandable Modality AlignmentabstractThis paper presents Babel, the expandable modality alignment model, specially designed for multi-modal sensing. While there has been considerable work on multi-modality alignment, they all struggle to effectively incorporate multiple sensing modalities due to the data scarcity constraints. How to utilize multi-modal data with partial pairings in sensing remains an unresolved challenge. Shenghong Dai, Shiqi Jiang 0002, Yifan Yang 0004, Ting Cao 0003, Mo Li 0001, Suman Banerjee 0001, Lili Qiu |
SenSys | 2 |
| 2025 | Efficient and Adaptive Diffusion Model Inference Through Lookup Table on Mobile DevicesabstractDiffusion models have revolutionized image synthesis applications. Many studies focus on using approximate computation such as model quantization to reduce inference costs on mobile devices. However, due to their extensive model parameters and autoregressive inference fashion, the overhead of diffusion models remains high, which is challenging for mobile devices to handle. To reduce the inference overhead of diffusion models on mobile devices, we proposeLUT-Diff, an algorithm-system co-design specifically tailored for mobile device diffusion model inference optimization.LUT-Diffoptimizes using lookup tables and can efficiently generate a series of lookup table candidates for diffusion models without end-to-end training. During inference,LUT-Diffadaptively selects the best inference strategy based on the application/user's latency budget. Additionally,LUT-Diffincludes a parallel inference engine that rapidly completes model inference through CPU-GPU co-scheduling. Extensive experiments demonstrate thatLUT-Diffcan generate images comparable to the original model, with an up to 0.012 MSE in generated images.LUT-Diffcan also achieve up to 9.1× inference acceleration and reduce the inference memory footprint by up to 70.9% compared to baseline methods. Moreover,LUT-Diffcan save at least 3281× the learning cost of lookup tables. Qipeng Wang 0001, Shiqi Jiang 0002, Yifan Yang 0004, Ruiqi Liu 0001, Yuanchun Li 0003, Ting Cao 0003, Xuanzhe Liu |
IEEE Trans. Mob. Comput. | 2 |
| 2025 | AdaWiFi, Collaborative WiFi Sensing for Cross-Environment AdaptationabstractDeep learning (DL) based Wi-Fi sensing has witnessed great development in recent years. Although decent results have been achieved in certain scenarios, Wi-Fi based activity recognition is still difficult to deploy in real smart homes due to the limited cross-environment adaptability, i.e. a well-trained Wi-Fi sensing neural network in one environment is hard to adapt to other environments. To address this challenge, we proposeAdaWiFi, a DL-based Wi-Fi sensing framework that allows multiple Internet-of-Things (IoT) devices to collaborate and adapt to various environments effectively. The key innovation ofAdaWiFiincludes a collective sensing model architecture that utilizes complementary information between distinct devices and avoids the biased perception of individual sensors and an accompanying model adaptation technique that can transfer the sensing model to new environments with limited data. We evaluate our system on a public dataset and a custom dataset collected from three complex sensing environments. The results demonstrate thatAdaWiFiis able to achieve significantly better sensing adaptation effectiveness (e.g. 30% higher accuracy with one-shot adaptation) as compared with state-of-the-art baselines. Naiyu Zheng, Yuanchun Li 0003, Shiqi Jiang 0002, Yuanzhe Li 0001, Rongchun Yao, Chuchu Dong, Zhimeng Yin 0001, Yunxin Liu 0001 |
IEEE Trans. Mob. Comput. | 3 |
| 2025 | Anatomizing Deep Learning Inference in Web BrowsersabstractWeb applications have increasingly adopted Deep Learning (DL) through in-browser inference , wherein DL inference performs directly within Web browsers. The actual performance of in-browser inference and its impacts on the Quality of Experience ( QoE ) remain unexplored, and urgently require new QoE measurements beyond traditional ones, e.g., mainly focusing on page load time. To bridge this gap, we make the first comprehensive performance measurement of in-browser inference to date. Our approach proposes new metrics to measure in-browser inference: responsiveness, smoothness, and inference accuracy. Our extensive analysis involves 9 representative DL models across Web browsers of 50 popular PC devices and 20 mobile devices. The results reveal that in-browser inference exhibits a substantial latency gap, averaging 16.9 times slower on CPU and 4.9 times slower on GPU compared to native inference on PC devices. The gap on mobile CPU and mobile GPU is 15.8 times and 7.8 times, respectively. Furthermore, we identify contributing factors to such latency gap, including underutilized hardware instruction sets, inherent overhead in the runtime environment, resource contention within the browser, and inefficiencies in software libraries and GPU abstractions. Additionally, in-browser inference imposes significant memory demands, at times exceeding 334.6 times the size of the DL models themselves, partly attributable to suboptimal memory management. We also observe that in-browser inference leads to a significant 67.2% increase in the time it takes for GUI components to render within Web browsers, significantly affecting the overall user QoE of Web applications reliant on this technology. Qipeng Wang 0001, Shiqi Jiang 0002, Zhenpeng Chen 0001, Yuanchun Li 0003, Aoyu Li, Yun Ma 0002, Ting Cao 0003, Xuanzhe Liu |
ACM Trans. Softw. Eng. Methodol. | 2 |
| 2024 | AutoDroid: LLM-powered Task Automation in AndroidabstractMobile task automation is an attractive technique that aims to enable voice-based hands-free user interaction with smartphones. However, existing approaches suffer from poor scalability due to the limited language understanding ability and the non-trivial manual efforts required from developers or endusers. The recent advance of large language models (LLMs) in language understanding and reasoning inspires us to rethink the problem from a model-centric perspective, where task preparation, comprehension, and execution are handled by a unified language model. In this work, we introduce AutoDroid, a mobile task automation system capable of handling arbitrary tasks on any Android application without manual efforts. The key insight is to combine the commonsense knowledge of LLMs and domain-specific knowledge of apps through automated dynamic analysis. The main components include a functionality-aware UI representation method that bridges the UI with the LLM, exploration-based memory injection techniques that augment the app-specific domain knowledge of LLM, and a multi-granularity query optimization module that reduces the cost of model inference. We integrate AutoDroid with off-the-shelf LLMs including online GPT-4/GPT-3.5 and on-device Vicuna, and evaluate its performance on a new benchmark for memory-augmented Android task automation with 158 common tasks. The results demonstrated that AutoDroid is able to precisely generate actions with an accuracy of 90.9%, and complete tasks with a success rate of 71.3%, outperforming the GPT-4-powered baselines by 36.4% and 39.7%. Hao Wen 0004, Yuanchun Li 0003, Guohong Liu 0002, Shanhui Zhao, Toby Jia-Jun Li, Shiqi Jiang 0002, Yunhao Liu 0001, Yunxin Liu 0001 |
MobiCom | 7 |
| 2024 | Empowering In-Browser Deep Learning Inference on Edge Through Just-In-Time Kernel OptimizationabstractWeb is increasingly becoming the primary platform to deliver AI services onto edge devices, making in-browser deep learning (DL) inference more prominent. Nevertheless, the heterogeneity of edge devices, combined with the underdeveloped state of Web hardware acceleration practices, hinders current in-browser inference from achieving its full performance potential on target devices. Fucheng Jia, Shiqi Jiang 0002, Ting Cao 0003, Tianrui Xia, Yuanchun Li 0003, Qipeng Wang 0001, Ju Ren 0001, Yunxin Liu 0001, Lili Qiu, Mao Yang 0004 |
MobiSys | 2 |
| 2024 | Large-scale Video Analytics with Cloud-Edge Collaborative Continuous LearningabstractDeep learning–based video analytics demands high network bandwidth to ferry the large volume of data when deployed on the cloud. When incorporated at the edge side, only lightweight deep neural network (DNN) models are affordable due to computational constraint. In this article, a cloud–edge collaborative architecture is proposed combining edge-based inference with cloud-assisted continuous learning. Lightweight DNN models are maintained at the edge servers and continuously retrained with a more comprehensive model on the cloud to achieve high video analytics performance while reducing the amount of data transmitted between edge servers and the cloud. The proposed design faces the challenge of constraints of both computation resources at the edge servers and network bandwidth of the edge–cloud links. An accuracy gradient-based resource allocation algorithm is proposed to allocate the limited computation and network resources across different video streams to achieve the maximum overall performance. A prototype system is implemented and experiment results demonstrate the effectiveness of our system with up to 28.6% absolute mean average precision gain compared with alternative designs. Ya Nan, Shiqi Jiang 0002, Mo Li 0001 |
ACM Trans. Sens. Networks | 2 |
| 2023 | AdaptiveNet: Post-deployment Neural Architecture Adaptation for Diverse Edge EnvironmentsabstractDeep learning models are increasingly deployed to edge devices for real-time applications. To ensure stable service quality across diverse edge environments, it is highly desirable to generate tailored model architectures for different conditions. However, conventional pre-deployment model generation approaches are not satisfactory due to the difficulty of handling the diversity of edge environments and the demand for edge information. In this paper, we propose to adapt the model architecture after deployment in the target environment, where the model quality can be precisely measured and private edge data can be retained. To achieve efficient and effective edge model generation, we introduce a pretraining-assisted on-cloud model elastification method and an edge-friendly on-device architecture search method. Model elastification generates a high-quality search space of model architectures with the guidance of a developer-specified oracle model. Each subnet in the space is a valid model with different environment affinity, and each device efficiently finds and maintains the most suitable subnet based on a series of edge-tailored optimizations. Extensive experiments on various edge devices demonstrate that our approach is able to achieve significantly better accuracy-latency tradeoffs (e.g. 46.74% higher on average accuracy with a 60% latency budget) than strong baselines with minimal overhead (13 GPU hours in the cloud and 2 minutes on the edge server). Hao Wen 0004, Yuanchun Li 0003, Zunshuai Zhang, Shiqi Jiang 0002, Xiaozhou Ye, Ye Ouyang, Yunxin Liu 0001 |
MobiCom | 4 |
| 2023 | NN-Stretch: Automatic Neural Network Branching for Parallel Inference on Heterogeneous Multi-ProcessorsabstractMobile devices are increasingly equipped with heterogeneous multiprocessors, e.g., CPU + GPU + DSP. Yet existing Neural Network (NN) inference fails to fully utilize the computing power of the heterogeneous multi-processors due to the sequential structures of NN models. Towards this end, this paper proposes NN-Stretch, a new model adaption strategy, as well as the supporting system. It automatically branches a given model according to the processor architecture characteristics. Compared to other popular model adaption techniques such as model pruning that often sacrifices accuracy, NN-Stretch accelerates inference while preserving accuracy. Jianyu Wei, Ting Cao 0003, Shijie Cao, Shiqi Jiang 0002, Shaowei Fu, Mao Yang 0004, Yanyong Zhang, Yunxin Liu 0001 |
MobiSys | 4 |
| 2022 | CoDL: efficient CPU-GPU co-execution for deep learning inference on mobile devicesabstractConcurrent inference execution on heterogeneous processors is critical to improve the performance of increasingly heavy deep learning (DL) models. However, available inference frameworks can only use one processor at a time, or hardly achieve speedup by concurrent execution compared to using one processor. This is due to the challenges to 1) reduce data sharing overhead, and 2) properly partition each operator between processors. Fucheng Jia, Ting Cao 0003, Shiqi Jiang 0002, Yunxin Liu 0001, Ju Ren 0001, Yaoxue Zhang |
MobiSys | 4 |
| 2022 | Turbo: Opportunistic Enhancement for Edge Video AnalyticsabstractEdge computing is being widely used for video analytics. To alleviate the inherent tension between accuracy and cost, various video analytics pipelines have been proposed to optimize the usage of GPU on edge nodes. Nonetheless, we find that GPU compute resources provisioned for edge nodes are commonly under-utilized due to video content variations, subsampling and filtering at different places of a video analytics pipeline. As opposed to model and pipeline optimization, in this work, we study the problem of opportunistic data enhancement using the non-deterministic and fragmented idle GPU resources. In specific, we propose a task-specific discrimination and enhancement module, and a model-aware adversarial training mechanism, providing a way to exploit idle resources to identify and transform pipeline-specific, low-quality images in an accurate and efficient manner. A multi-exit enhancement model structure and a resource-aware scheduler is further developed to make online enhancement decisions and fine-grained inference execution under latency and GPU resource constraints. Experiments across multiple video analytics pipelines and datasets reveal that our system boosts DNN object detection accuracy by 7.27 -- 11.34% by judiciously allocating 15.81 -- 37.67% idle resources on frames that tend to yield greater marginal benefits from enhancement. Yan Lu 0006, Shiqi Jiang 0002, Ting Cao 0003, Yuanchao Shu |
SenSys | 2 |
| 2021 | Flexible high-resolution object detection on edge devices with tunable latencyabstractObject detection is a fundamental building block of video analytics applications. While Neural Networks (NNs)-based object detection models have shown excellent accuracy on benchmark datasets, they are not well positioned for high-resolution images inference on resource-constrained edge devices. Common approaches, including down-sampling inputs and scaling up neural networks, fall short of adapting to video content changes and various latency requirements. This paper presents Remix, a flexible framework for high-resolution object detection on edge devices. Remix takes as input a latency budget, and come up with an image partition and model execution plan which runs off-the-shelf neural networks on non-uniformly partitioned image blocks. As a result, it maximizes the overall detection accuracy by allocating various amount of compute power onto different areas of an image. We evaluate Remix on public dataset as well as real-world videos collected by ourselves. Experimental results show that Remix can either improve the detection accuracy by 18%-120% for a given latency budget, or achieve up to 8.1× inference speedup with accuracy on par with the state-of-the-art NNs. Shiqi Jiang 0002, Yuanchun Li 0003, Yuanchao Shu, Yunxin Liu 0001 |
MobiCom | 1 |
| 2019 | Memento: An Emotion-driven Lifelogging System with WearablesabstractDue to the increasing popularity of mobile devices, the usage of lifelogging has dramatically expanded. People collect their daily memorial moments and share with friends on the social network, which is an emerging lifestyle. We see great potential of lifelogging applications along with rapid recent growth of the wearables market, where more sensors are introduced to wearables, i.e., electroencephalogram (EEG) sensors, that can further sense the user’s mental activities, e.g., emotions. In this article, we present the design and implementation of Memento, an emotion-driven lifelogging system on wearables. Memento integrates EEG sensors with smart glasses. Since memorable moments usually coincides with the user’s emotional changes, Memento leverages the knowledge from the brain-computer-interface domain to analyze the EEG signals to infer emotions and automatically launch lifelogging based on that. Towards building Memento on Commercial off-the-shelf wearable devices, we study EEG signals in mobility cases and propose a multiple sensor fusion based approach to estimate signal quality. We present a customized two-phase emotion recognition architecture, considering both the affordability and efficiency of wearable-class devices. We also discuss the optimization framework to automatically choose and configure the suitable lifelogging method (video, audio, or image) by analyzing the environment and system context. Finally, our experimental evaluation shows that Memento is responsive, efficient, and user-friendly on wearables. Shiqi Jiang 0002, Zhenjiang Li 0001, Mo Li 0001 |
ACM Trans. Sens. Networks | 1 |
| 2017 | Memento: An Emotion Driven Lifelogging System with WearablesabstractDue to the increasing popularity of mobile devices, the usage of lifelogging has been dramatically expanded. People collect their daily memorial moments and share with friends on the social network, which has been an emerging lifestyle. We see great potential of lifelogging applications along with rapid growth of recent wearable market, where more sensors are introduced to wearables, i.e., electroencephalogram (EEG) sensors, that can further sense the user's mental activities, e.g., emotions. In this paper, we present the design and implementation of Memento, an emotion driven lifelogging system on wearables. Memento integrates EEG sensors with smart glasses. Since memorable moments usually coincides with the user's emotional changes, Memento leverages the knowledge from the brain-computer-interface (BCI) domain to analyze the EEG signals to infer emotions and automatically launch lifelogging based on that. Towards building Memento on COTS wearable devices, we study EEG signals in mobility cases and propose a multiple sensor fusion based approach to estimate signal quality. We also present a customized two-phase emotion recognition architecture, considering both the affordability and efficiency of wearable-class devices. Our experimental evaluation shows that Memento is responsive, efficient and user-friendly on wearables. Shiqi Jiang 0002, Zhenjiang Li 0001, Mo Li 0001 |
ICCCN | 1 |
| 2017 | A Participatory Urban Traffic Monitoring System: The Power of Bus RidersabstractThis paper presents a participatory sensing-based urban traffic monitoring system. Different from existing works that heavily rely on intrusive sensing or full cooperation from probe vehicles, our system exploits the power of participatory sensing and crowdsources the traffic sensing tasks to bus riders' mobile phones. The bus riders are information source providers and, meanwhile, major consumers of the final traffic output. The system takes public buses as dummy probes to detect road traffic conditions, and collects the minimum set of cellular data together with some lightweight sensing hints from the bus riders' mobile phones. Based on the crowdsourced data from participants, the system recovers the bus travel information and further derives the instant traffic conditions of roads covered by bus routes. The real-world experiments with a prototype implementation demonstrate the feasibility of our system, which achieves accurate and fine-grained traffic estimation with modest sensing and computation overhead at the crowd. Zhidan Liu 0001, Shiqi Jiang 0002, Mo Li 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2015 | Urban Traffic Monitoring with the Help of Bus RidersabstractReal-time urban traffic conditions are critical to wide populations in the city and serve the needs of many transportation dependent applications. This paper presents our experience of building a participatory urban traffic monitoring system that exploits the power of bus riders' mobile phones. The system takes lightweight sensor hints and collects minimum set of cellular data from the bus riders' mobile phones. Based on such a participatory sensing framework, the system turns buses into dummy probes, monitors their travel statuses, and derives the instant traffic map of the city. Unlike previous works that rely on intrusive detection or full cooperation from "probe vehicles", our approach resorts to the crowd-participation of ordinary bus riders, who are the information source providers and major consumers of the final traffic output. The experiment results demonstrate the feasibility of such an approach achieving fine-grained traffic estimation with modest sensing and computation overhead at the crowd. Shiqi Jiang 0002, Mo Li 0001 |
ICDCS | 2 |
| 2014 | Demo: instant phone attitude estimation and its applicationsabstractThe phone attitude is an essential input to many smartphone applications. Based on in-depth understanding of the nature of the MEMS gyroscope and other IMU sensors, we propose A3 - an accurate and automatic attitude detector for commodity smartphones. In the demo, we show the performance of our attitude tracking algorithm and its usability in attitude-based mobile applications. Weiming Chan, Shiqi Jiang 0002, Jiajue Ou, Mo Li 0001, Guobin Shen |
MobiCom | 3 |