VLDB 2026 Research / reviewers in the wild / expert
Mu Yuan
dblp:08/9064
· DBLP profile ↗
41ranked-venue papers
11as first author
36since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 23 · 6 first-author · 21 since 2021Databases, data management, data science and information retrieval · 6 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Theory of computation · 4 · 3 since 2021Security and privacy · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TopFGL: A Topology-Aware and Distributionagnostic Federated Learning Framework Tackling Topological Heterogeneity on Graph Data
Junyang Wang 0004, Lan Zhang 0002, Yihang Cheng 0002, Mu Yuan, Tian Wang 0001, Zhihui Fu |
ICDE | 4 |
| 2026 | BeeKeeper: Securing Cross-Technology Communication via Channel-Aware Dual-Binding
Weizheng Wang 0001, Qipeng Xie, Mu Yuan, Qingqing Ye 0001, Kaishun Wu, Haibo Hu 0001 |
INFOCOM | 3 |
| 2026 | Venus: An Efficient Edge Memory-and-Retrieval System for VLM-based Online Video Understanding
Shengyuan Ye, Bei Ouyang, Tianyi Qian, Liekang Zeng, Mu Yuan, Xiaowen Chu 0001, Weijie Hong, Xu Chen 0004 |
INFOCOM | 5 |
| 2026 | A Large-Scale Multimodal Dataset and Benchmarks for Human Activity Scene Understanding and ReasoningabstractMultimodal human action recognition (HAR) utilizes complementary data for activity classification. Built on traditional HAR tasks, recent advances in Large Language Models (LLMs) enable detailed descriptions and causal reasoning of human actions, advancing new tasks of human action understanding (HAU) and human action reasoning (HARn). However, most LLMs, especially multimodal Large Vision-Language Models (LVLMs), struggle with modalities other than RGB images, like depth, IMU, ormmWave, due to a lack of large-scale datasets in these task domains. Existing HAR datasets provide only coarse-grained annotations, in-sufficient for depicting the detailed action dynamics required in HAU and HARn tasks. Simply combining annotations and generating captions with LLMs often lacks necessary logical and spatiotemporal consistency. In this paper, we introduce CUHK-X, a large-scale multi-modal dataset and benchmarks for HAR, HAU, and HARn. It includes 64,267 samples of 40 actions performed by 30 participants across two indoor environments, covering diverse daily scenarios. To address the challenge of spatiotemporal inconsistencies in captions, we propose a prompt-based scene creation method that leverages LLMs to generate logically connected activity sequences. CUHK-X also includes three benchmarks with six tasks to evaluate state-of-the-art models. Experimental results show average accuracies of 76.52% for HAR, 40.76% for HAU, and 70.25% for HARn. This large-scale multimodal dataset aims to empower the research community to apply, develop, and adapt data-intensive learning techniques for a wide range of human activity-related tasks. Siyang Jiang, Mu Yuan, Bufang Yang, Lilin Xu, Yang Li 0147, Yuting He 0006, Liran Dong, Wenrui Lu, Zhenyu Yan 0002, Xiaofan Jiang 0001, Wei Gao 0006, Hongkai Chen 0001, Guoliang Xing |
MobiSys | 2 |
| 2026 | STIP: Three-Party Privacy-Preserving and Lossless Inference for Large Transformers in Production
Mu Yuan, Lan Zhang 0002, Yihang Cheng 0002, Miaohui Song, Guoliang Xing, Xiang-Yang Li 0001 |
NDSS | 1 |
| 2026 | COACH: Adaptive Robust Human-Robot Collaboration for Efficient Smart ManufacturingabstractModern smart manufacturing pipelines have pervasively collaborated human workers, mobile robots, and industrial Internet of Things (IIoT) in shared workspaces for versatile production tasks. Despite the promising capacity of individual entities, the performance of these IIoT systems largely relies on pipeline coordination, i.e., task dispatching between humans and robots, which is particularly challenging under heterogeneous physical constraints and complex environmental uncertainties. Nonetheless, existing works either rely on traditional operation frameworks that lack scalability for large-scale complex production, or propose customized solutions for fixed agent models, overlooking the evolving nature of IIoT environments. To address these limitations, this paper proposes COACH, a human-robot collaborative manufacturing system that enables robust constraint-aware coordination across humans, robots, and IIoT. Specifically, COACH designs a scalable contextual encoder to represent the evolving relationships among human and robot agents in dynamic heterogeneous graphs. With that, a novel experience-driven task dispatcher is developed, enabling both high-performance and computation-efficient policy generation concerning the status of IIoT. To accommodate changing human fatigue and pipeline scales, COACH further develops a curriculum-enhanced reinforcement learning module for efficient dispatcher adaptation. Extensive evaluations using both synthetic testbeds and real-world manufacturing datasets demonstrate that COACH improves the feasible ratio of manufacturing pipelines by up to 27.4% and achieves up to 13.9% improvement in time efficiency compared to competing baselines across diverse job scales and environmental settings. Hui Wang 0011, Liekang Zeng, Zhiwen Yu 0001, Yao Zhang 0005, Di Duan, Mu Yuan, Bin Guo 0001, Guoliang Xing |
SenSys | 6 |
| 2026 | A Generalized χn-FunctionabstractThe mappingχnfrom Fn2to itself defined byy=χn(x) withyi=xi+xi+2(1 +xi+1), where the indices are computed modulo n, has been widely studied for its applications in lightweight cryptography. However,χnis bijective on Fn2only whennis odd, restricting its use to odd-dimensional vector spaces over F2. To address this limitation, we introduce and analyze the generalized mappingχn,mdefined byy=χn,m(x) withyi=xi+xi+m(xi+m−1+ 1)(xi+m−2+ 1) · · · (xi+1+ 1), wheremis a fixed integer withm∤n. To investigate such mappings, we further generalizeχn,mto θm,k, where θm,kis given byyi=xi+mkПmk−1j=1, m∤j(xi+j+ 1) , fori∈ {0, 1, . . . ,n− 1}. We prove that these mappings generate an abelian group isomorphic to the group of units in F2[z]/(z⌊n/m⌋+1). This structural insight enables us to construct a broad class of permutations over Fn2for any positive integern, along with their inverses. We rigorously analyze algebraic properties of these mappings, including their iterations, fixed points, and cycle structures. Additionally, we provide a comprehensive database of the cryptographic properties for iterates ofχn,mfor small values ofnandm. Finally, we conduct a comparative security and implementation cost analysis amongχn,m,χn,χχn(EUROCRYPT 2025 [3]) and their variants, and prove Conjecture 1 proposed in [3] as a by-product of our study. Our results lead to generalizations ofχn, providing alternatives toχnandχχn. Mu Yuan, Dabin Zheng, Siwei Sun, Shun Li 0004 |
IEEE Trans. Inf. Theory | 2 |
| 2026 | Gproxy: Communication-Efficient Federated Graph Learning With Efficient Adaptive ProxyingabstractFederated graph learning (FGL) enables multiple participants with distributed but connected graph data to collaboratively train a model in a privacy-preserving way. However, the high communication cost hinders the adoption of FGL in many resource-limited or delay-sensitive applications. In this work, we focus on reducing the communication cost incurred by the transmission of neighborhood information in FGL. We propose to search for local proxies that can play a substitute role as the external neighbors and develop a novel federated graph learning framework namedGproxy.Gproxyutilizes representation similarity and class correlation to select local proxies for external neighbors. Additionally, we propose to dynamically adjust the proxy strategy according to the changing representation of nodes during the iterative training process. We also design a proxy cache to accelerate the search process by reusing proxy search outcomes for similar external neighbors. Furthermore, we provide a theoretical analysis and show that using a proxy node has a similar influence on training when it is sufficiently similar to the external one. Extensive evaluations show thatGproxysignificantly reduces communication cost while maintaining model performance compared to strong baselines. Junyang Wang 0004, Lan Zhang 0002, Mu Yuan, Yihang Cheng 0002, Yunhao Yao, Zhonghao Hu |
IEEE Trans. Mob. Comput. | 3 |
| 2025 | A-VL: Adaptive Attention for Large Vision-Language ModelsabstractThe Large Vision-Language Model (LVLM) integrates computer vision and natural language processing techniques, offering substantial application potential. However, these models demand extensive resources during inference. Adaptive attention techniques can dynamically reduce computational redundancy and thus improve efficiency. Although current adaptive attention methods significantly reduce the memory requirements of Transformer-based language models, they are not tailored for LVLMs. We observe that LVLMs generate responses from both remote image tokens and local text tokens, and different modalities have different attention patterns. This observation inspires us to manage the attention for each modality separately. Specifically, for visual input, we store the cache of potentially useful information but only compute the most critical parts. For language input, we care more about local information. Based on our observation and analysis of vision-language attention patterns, we develop A-VL, a plug-and-play adaptive attention tailored for LVLM inference. Extensive evaluations on three vision-language tasks and five datasets show the effectiveness of our designs. Our approach A-VL outperforms existing adaptive attention methods in reducing memory usage and computational load without compromising performance. Junyang Zhang 0001, Mu Yuan, Ruiguang Zhong, Puhan Luo, Huiyou Zhan, Ningkang Zhang, Chengchen Hu, Xiang-Yang Li 0001 |
AAAI | 2 |
| 2025 | ProxySampler: Proxy Informativeness Estimation for Efficient Data Selection in Active LearningabstractLarge-scale data analysis services require efficient periodic model updates to adapt to the possibly changing data distributions. Manually labeling all available samples for task model updates is infeasible for a large sample scale. Active learning technique is proposed to iteratively select subsets of the most informative samples for labeling. From our experience of applying active learning in a real-world video analysis system, we identify a previously overlooked bottleneck of time cost: data selection. Existing active learning methods select data by estimating informativeness (e.g., output confidence) over all unlabeled samples in each iteration. This data selection process can take up to 42% of the time cost of end-to-end model updates in our system (totals include the time for manual labeling, data selection, and model updates.). To address the time cost bottleneck caused by data selection, we propose a new idea: proxy informativeness estimation. We start with modeling the time cost of data selection, from which we identify three key factors: unit estimation cost, the number of samples for estimation, and the number of iteration rounds. The influence of the first two factors increases cumulatively with the number of iteration rounds. Correspondingly, we design a proxy estimator and a sample pooling method, respectively. Our proxy estimator is a lightweight neural network for direct informativeness estimation to replace the role of the high-cost task model, thus reducing the unit cost. And, our sample pooling method leverages historical estimation results to narrow the scope of sample candidates. Based on the above design, we develop ProxySampler, which can be integrated with various active learning approaches as a plug-in. Experimental results show that integrating ProxySampler with state-of-the-art active learning methods can reduce the time cost by 53.6-83.3% (a 2.15-6.01x speedup) when achieving the same accuracy. Miaohui Song, Lan Zhang 0002, Mu Yuan, Yijun Liu 0003 |
CIKM | 3 |
| 2025 | Myo-Trainer: A Vision-based Muscle-Aware Motion Feedback System for In-Home Resistance TrainingabstractIn-home resistance training (RT) is a convenient and effective way to maintain health and well-being. However, incorrect exercise execution can result in unintended muscle engagement and an increased risk of injury. Without access to professional coaching, an accurate muscle-aware motion feedback system becomes essential for safe and effective training. However, existing visual language models (VLMs) struggle to provide accurate and effective muscle-aware movement guidance due to their limited understanding of RT motion and the absence of related expert knowledge. In this work, we introduce Myo-Trainer, the first vision-based muscle-aware motion feedback system that uses explicit muscle-aware motion analysis and domain-specific expert knowledge to provide corrective guidance on muscle engagement and movement execution. Also, we propose a novel DAGCN-Former network that integrates both spatial and temporal modeling capabilities to capture the complex dynamics of human RT motion. Experiments involving 26 subjects and 1000+ minutes of RT demonstrate that Myo-Trainer improves the accuracy of motion analysis by 17.22%, achieves a 2.5x reduced inference latency and a BertScore of 85.88% of generated feedback compared to those provided by experienced certified trainers, outperforming existing solutions. Additionally, Myo-Trainer received higher satisfaction ratings from participants compared to other AI trainers and video tutorials, highlighting its potential for real-world applications. Yuting He 0006, Xinyan Wang 0003, Mu Yuan, Bufang Yang, Siyang Jiang, Yihua Huang 0002, Doris Sau-Fung Yu, Guoliang Xing, Hongkai Chen 0001 |
MobiCom | 3 |
| 2025 | Privacy-Preserving LLM Agent for Multi-modal Health Monitoring
Qipeng Xie, Jiafei Wu, Zhuotao Lian, Mu Yuan, Xian Shuai, Weizheng Wang 0001, Yuan Haoyi, Haibo Hu 0001, Kaishun Wu |
ProvSec | 5 |
| 2025 | Grape: Efficient Spatiotemporal Prediction Services with Stale Sensing StreamsabstractEmerging cyber-physical systems have embraced a large number of IoT devices spanning geo-distributed, which generate and consume massive volumes of data continuously. Accurate and timely spatiotemporal predictions (STP) over these streaming sensor data are critical and, in growing demand, ubiquitous across various edge scenarios such as traffic flow forecasting. Towards that, recent advanced systems have developed sophisticated optimizations among STP pipelines, aiming at optimal prediction performance. However, based on our empirical studies in real-world settings, we identify a previously overlooked bottleneck of end-to-end STP performance: data staleness. To mitigate this issue, in this work, we investigate a new task, namely stream interception, which deliberately terminates the acceptance of incoming sensor data and anticipates model execution with imputed missing features. We propose a novel dynamic interception strategy to determine the time slot to exit waiting and present Grape, an STP system that implements it with practical system designs. Extensive evaluations on real-world traces show that Grape can strike a superior tradeoff between prediction accuracy and serving latency, achieving 1.69-1.90× speedup against traditional all-waiting baselines across various STP services with high prediction accuracy on par with offline optimal cases. Liekang Zeng, Shengyuan Ye, Mu Yuan, Di Duan, Xu Chen 0004, Guoliang Xing |
RTSS | 4 |
| 2025 | Argus: Multi-View Egocentric Human Mesh Reconstruction Based on Stripped-Down Wearable mmWave Add-onabstractIn this paper, we propose Argus, a wearable add-on system based on stripped-down (i.e., compact, lightweight, low-power, limited-capability) mmWave radars. It is the first to achieve egocentric human mesh reconstruction in a multi-view manner. Compared with conventional frontal-view mmWave sensing solutions, it addresses several pain points, such as restricted sensing range, occlusion, and the multipath effect caused by surroundings. To overcome the limited capabilities of the stripped-down mmWave radars (with only one transmit antenna and three receive antennas), we tackle three main challenges and propose a holistic solution, including tailored hardware design, sophisticated signal processing, and a deep neural network optimized for high-dimensional complex point clouds. Extensive evaluation shows that Argus achieves performance comparable to traditional solutions based on high-capability mmWave radars, with an average vertex error of 6.5 cm, solely using stripped-down radars deployed in a multi-view configuration. It presents robustness and practicality across conditions, such as with unseen users and different host devices. Di Duan, Shengzhe Lyu, Mu Yuan, Hongfei Xue, Tianxing Li 0001, Weitao Xu, Kaishun Wu, Guoliang Xing |
SenSys | 3 |
| 2025 | SCX: Stateless KV-Cache Encoding for Cloud-Scale Confidential Transformer ServingabstractTransformer models have revolutionized fields like natural language processing and computer vision but face privacy concerns in sensitive applications such as medical diagnostics. Existing confidential serving methods, including cryptography-based, memory isolation-based, and access control-based, offer trade-offs between privacy and efficiency but often struggle with high latency or hardware dependencies. This work proposes stateless KV-cache encoding (SCX), a novel framework that encodes the intermediate key-value cache during Transformer inference using user-controlled keys. SCX ensures that the cloud can neither recover the input nor independently complete the next token prediction, effectively preserving privacy. By introducing efficient encoding and decoding schemes, SCX addresses communication complexity and attack vulnerabilities while ensuring zero loss of inference quality. Experiments on large Transformer models demonstrate that SCX achieves lower latency (e.g., 36ms for LLaMA-7B), outperforming state-of-the-art cryptography and memory isolation methods by orders of magnitude. Moreover, SCX can complementarily work with advanced KV-cache management techniques to further enhance KV-cache communication efficiency by 85%, marking a significant step toward practical, privacy-preserving large Transformer serving. Mu Yuan, Lan Zhang 0002, Liekang Zeng, Siyang Jiang, Bufang Yang, Di Duan, Guoliang Xing |
SIGCOMM | 1 |
| 2025 | Mitigating Tail Latency for On-Device Inference With Load-Balanced Heterogeneous ModelsabstractServing machine learning models on edge, mobile, and embedded devices places stringent requirements on inference latency. From operating a real enterprise service, we observed that even a fully optimized model could lead to severe violations of latency objectives when the load surges. A straightforward and mature approach is to auto-scale multiple models to balance the load. However, unlike cloud clusters, edge or mobile devices usually cannot afford to deploy multiple model replicas. Therefore, in this paper, we explore a new idea: in addition to the original model, we deploy one (or more) heterogeneous model(s) with much smaller resource overhead on the device, and perform load balancing among all models. We overcame the technical challenges posed by performance dynamics and developed InferRouter based on queuing theory. We implement and evaluate InferRouter on three real on-device inference systems, covering mobile sensing, video analytics, and natural language processing applications. Experimental results show that compared with strong baselines, InferRouter can decrease 85.2% P99 latency (5.8x faster) and improve 5.9% accuracy on the mobile workload. For a traffic video analytics task, InferRouter achieves 55.1% higher accuracy with zero deadline misses. InferRouter also shows its advantages in saving resources compared with auto-scaling and offloading approaches. Mu Yuan, Lan Zhang 0002, Di Duan, Liekang Zeng, Miaohui Song, Zichong Li, Guoliang Xing, Xiang-Yang Li 0001 |
IEEE Trans. Mob. Comput. | 1 |
| 2025 | TrafficDiary: User Attribute Inference Based on Smart Home Traffic TracesabstractSmart home technology has found wide-ranging applications in daily life, from enhancing energy efficiency to simplifying daily tasks and providing greater convenience. However, recent works have found that smart home devices are vulnerable to passive network observers (i.e., adversaries). While adversaries have demonstrated the ability to infer device events (e.g., whether a lamp is turned on) from the encrypted smart home traffic, we believe this only represents a less critical aspect of smart home privacy risks. Further analysis of demographic attributes presents greater risks to user privacy. Besides, from our deployment experience of real-world smart homes, we found that existing event inference methods can be greatly interfered with by event-unrelated traffic. Experiments show that this interference can result in up to a 10% drop in inference accuracy. Furthermore, it is challenging to infer finer-grained demographic attributes, due to the insufficient accuracy of event inference. Therefore, in this work, we propose a novel event inference model extracting multi-dimensional features that reduces the interference of event-unrelated traffic by analyzing packet length distribution and statistical properties. In addition, we design a dual-channel neural network to extract spatial and temporal relationships among triggered events to infer demographic attributes of smart home users, such as age group and career stage. Combining the above designs, we present TrafficDiary, the first user attribute inference approach based on smart home traffic traces. We prototype TrafficDiary and evaluate it in real-world smart homes. Experimental results show that TrafficDiary achieves 98.68% accuracy with a zero false positive rate in event inference and a high level of accuracy in user attribute inference, even when 16, 362 groups of event-unrelated traffic exist. TrafficDiary also performs well in terms of efficiency, with an inference latency of only 1.82 ms on a Raspberry Pi 4B device. Yunhao Yao, Jiahui Hou, Mu Yuan, Zhengyuan Xu, Xiang-Yang Li 0001 |
ACM Trans. Internet Techn. | 3 |
| 2025 | ChannelZip: SLO-Aware Channel Compression for Task-Adaptive Model Serving on IoT DevicesabstractDeploying deep neural networks (DNNs) on IoT devices for model serving is a promising solution for intelligent applications with high real-time requirements and bandwidth sensitivity. To cope with the prohibitive computation and storage overheads of modern DNNs, great efforts have been devoted to the model compression technique. Most existing model compression approaches focus on minimizing the model size and maximizing the average accuracy on all the inference tasks. However, real-world IoT tasks have various service-level objectives (SLOs). Models compressed by existing methods struggle to simultaneously meet SLOs in multiple dimensions, such as latency and accuracy. In this work, we study model compression with a joint consideration of SLO awareness and task adaptation. Through our extensive experience with model compression across various IoT tasks, we observe that the importance of individual channels in contributing to accuracy is heavily influenced by task-specific data distribution. Therefore, we design a channel Shapley algorithm to estimate the importance of individual channels in DNNs and propose a deep reinforcement learning based controller to incorporate SLOs into the compression objective. Integrating these designs, we propose and prototype ChannelZip, the first SLO-aware channel compression framework. Extensive evaluations on real IoT model serving systems show the effectiveness in task adaptation of ChannelZip. ChannelZip outperforms strong model compression baselines by 3.77% accuracy and achieves a 69% average parameter compression ratio. Real-world deployment on different IoT devices shows that ChannelZip meets all task SLOs and achieves up to 2.32 × inference speedup. Puhan Luo, Jiahui Hou, Haisheng Tan, Mu Yuan, Xiang-Yang Li 0001 |
ACM Trans. Sens. Networks | 4 |
| 2024 | GraphProxy: Communication-Efficient Federated Graph Learning with Adaptive ProxyabstractFederated graph learning (FGL) enables multiple participants with distributed but connected graph data to collaboratively train a model in a privacy-preserving way. However, the high communication cost hinder the adoption of FGL in many resource-limited or delay-sensitive applications. In this work, we focus on reducing the communication cost incurred by the transmission of neighborhood information in FGL. We propose to search for local proxies that can play a substitute role as the external neighbors, and develop a novel federated graph learning framework named GraphProxy. GraphProxy utilizes representation similarity and class correlation to select local proxies for external neighbors. And we propose to dynamically adjust the proxy strategy according to the changing representation of nodes during the iterative training process. We also perform a theoretical analysis and show that using a proxy node has a similar influence on training when it is sufficiently similar to the external one. Extensive evaluations show the effectiveness of our design, e.g., GraphProxy can achieve 8× communication efficiency with only 0.14% performance degradation. Junyang Wang 0004, Lan Zhang 0002, Mu Yuan, Yihang Cheng 0002 |
INFOCOM | 4 |
| 2024 | Demo: Myotrainer: Muscle-Aware Motion Analysis and Feedback System for In-Home Resistance TrainingabstractResistance training is widely incorporated in exercise programs, including in-home fitness and rehabilitation. However, improper motion patterns and muscle stimulation can undermine the safety of the subjects, making precise monitoring essential. Existing solutions primarily focus on correcting motion patterns with difficulties assessing muscle contraction levels. In this work, we introduce MyoTrainer, which provides muscle-aware motion descriptions and personalized feedback in natural language. Taking a person's exercise video as input, MyoTrainer first utilizes pose estimation models to capture motion sequences in real-time. A GCN-Former model has been developed for fine-grained motion analysis, which includes action recognition, incorrect movement pattern detection, and muscle contraction intensity estimation. Additionally, MyoTrainer integrates fitness and physiotherapeutic domain knowledge to deliver personalized, professional feedback. Extensive evaluations show that our system outperforms existing solutions in all recognition tasks and a survey indicates 88.9% of users find the generated feedback to be beneficial. Yuting He 0006, Xinyan Wang 0003, Mu Yuan, Di Duan, Doris Sau-Fung Yu, Guoliang Xing, Hongkai Chen 0001 |
SenSys | 3 |
| 2024 | F2Zip: Finetuning-Free Model Compression for Scenario-Adaptive Embedded VisionabstractWith the development of the Internet of Things and artificial intelligence, the deployment and inference of intelligent models have gradually raised concerns. To reduce the huge computation and storage overhead of modern deep neural networks, many studies use model pruning techniques to reduce the model size and computational cost. However, existing pruning techniques usually require model fine-tuning, which incurs high additional overhead, making them difficult to apply to real-world scenarios. In this work, we focus on vision model compression and present F2Zip, a scenario-adaptive finetuning-free pruning framework for embedded devices. First, we propose a scenario complexity measurement that quantifies scenario changes with pixel-level entropy. By analyzing the scenario complexity, F2Zip adaptively evaluates the importance of different channels and layers of the model using only a small amount (tens) of unlabeled data. Then we design a multi-constraint knapsack solver to prune scenario-unrelated redundant channels. We implemented and deployed F2Zip in surveillance scenarios and tested different models on videos collected from both public and real-world sources. Experimental results show that F2Zip is free of model fine-tuning in various scenarios. F2Zip reduces the end-to-end deployment time by 89.8% and reduces energy cost by 79.5%, which shows that F2Zip is computationally friendly for embedded devices. Without fine-tuning and any accuracy degradation, F2Zip achieves up to 50.2% parameter reduction, outperforming baseline methods by 35.1%. Puhan Luo, Jiahui Hou, Mu Yuan, Yunhao Yao, Xiang-Yang Li 0001 |
SenSys | 3 |
| 2024 | FusionFlow: Neural Fusion and Compression for Communication-Efficient Edge-Cloud Collaborative Computing
Ningkang Zhang, Mu Yuan, Xiang-Yang Li 0001 |
WASA (2) | 4 |
| 2024 | InFi: End-to-End Learning to Filter Input for Resource-Efficiency in Mobile-Centric InferenceabstractMobile-centric AI applications have high requirements for the resource-efficiency of model inference. Input filtering is a promising approach to eliminate redundancy so as to reduce the cost of inference. Previous efforts have tailored effective solutions for many applications, but left two essential questions unanswered: (1)theoretical filterability of an inference workloadto guide the application of input filtering techniques, thereby avoiding the trial-and-error cost for resource-constrained mobile applications; (2)robust discriminability of feature embeddingto allow input filtering to be widely effective for diverse inference tasks and input content. To answer them, we first formulate the input filtering problem and theoretically compare the hypothesis complexity of inference models and input filters to understand the optimization potential. Then we propose the first end-to-end learnable input filtering framework that covers most state-of-the-art methods and surpasses them in feature embedding with robust discriminability. We design and implementInFithat supports different input modalities and mobile-centric deployments. Comprehensive evaluations confirm our theoretical results and show thatInFioutperforms strong baselines in applicability, accuracy, and efficiency.InFican achieve 8.5× throughput and save 95% bandwidth, while keeping over 90% accuracy, for a video analytics application on mobile platforms. Mu Yuan, Lan Zhang 0002, Fengxiang He, Xueting Tong, Miaohui Song, Zhengyuan Xu, Xiang-Yang Li 0001 |
IEEE Trans. Mob. Comput. | 1 |
| 2024 | SecoInfer: Secure DNN End-Edge Collaborative Inference Framework Optimizing Privacy and LatencyabstractEnd-edge collaborative inference enhances computational efficiency by segmenting a deep neural network (DNN) model into two parts, executed across the end device and the edge node. However, existing collaborative inference strategies often involve transmitting original inputs from the end device to the edge node, resulting in significant risks of user detail leakage without requiring input reconstruction. Therefore, in this work, we present SecoInfer, a secure layer-level DNN end-edge collaborative inference framework. SecoInfer achieves joint optimization of data privacy and inference latency for DNN partition solutions that meet latency constraints, supported by three key designs. First, the privacy-aware DNN layer projection measurement quantifies the difficulty adversaries encounter in reconstructing the original input from the intermediate output of each layer. Then, the latency-privacy integrated structure modeling enables the direct calculation of the privacy measurement and inference latency for each partition solution from a list element or a directed acyclic graph (DAG) cut. Finally, the two-stage latency constraint adjustment scheme narrows down the search space of feasible partition solutions at the block level and fine-tunes the final one to meet the latency constraint based on layer depth. We prototype SecoInfer, utilizing a Raspberry Pi 4B as the end device and a server with an NVIDIA GeForce RTX 3060 GPU as the edge node. Experimental results demonstrate that under latency constraints of 20 ms, 33 ms, and 40 ms, SecoInfer reduces adversarial data reconstruction by 9.84%, 19.26%, and 25.18%, respectively, without any loss of task model accuracy. SecoInfer also enhances efficiency, reducing the time needed to determine optimal end-edge partition solutions on a Raspberry Pi 4B by 18.04%. Yunhao Yao, Jiahui Hou, Yihang Cheng 0002, Mu Yuan, Puhan Luo, Xiang-Yang Li 0001 |
ACM Trans. Sens. Networks | 5 |
| 2023 | Efficient Deep Ensemble Inference via Query Difficulty-dependent Task SchedulingabstractDeep ensemble learning has been widely adopted to boost accuracy through combing outputs from multiple deep models prepared for the same task. However, the extra computation and memory cost it entails could impose an unacceptably high deadline miss rate in latency-sensitive tasks. Conventional approaches, including ensemble selection, focus on accuracy while ignoring deadline constraints, and thus cannot smartly cope with bursty query traffic and queries with different hardness. This paper explores redundancy in deep ensemble model inference and presents Schemble, a query difficulty-dependent task scheduling framework. Schemble treats ensemble inference progress as multiple base model inference tasks and schedules tasks for queries based on their difficulty and queuing status. We evaluate Schemble on real-world datasets, considering intelligent Q&A system, video analysis and image retrieval as the running applications. Experimental results show that Schemble achieves a 5× lower deadline miss rate and improves the accuracy by 30.8% given deadline constraints. Zichong Li, Lan Zhang 0002, Mu Yuan, Miaohui Song, Qi Song 0004 |
ICDE | 3 |
| 2023 | PacketGame: Multi-Stream Packet Gating for Concurrent Video Inference at ScaleabstractThe resource efficiency of video analytics workloads is critical for large-scale deployments on edge nodes and cloud clusters. Recent advanced systems have benefited from techniques including video compression, frame filtering, and deep model acceleration. However, based on our year-long experience of operating a real-time video analytics system on more than 1000 cameras, we identified a previously overlooked bottleneck of end-to-end concurrency: video decoding. To support concurrent video inference at scale, in this work, we investigate a new task, named video packet gating, which selectively filters packets before running a decoder. We propose a novel multi-view embedding approach for video packets and present PacketGame that has both theoretical performance guarantee and practical system designs. Experiments on both public datasets and a real system show PacketGame saves 52.0--79.3% decoding costs and achieves 2.1--4.8× concurrency compared to original workloads. Comparisons with four state-of-the-art complementary methods show the superiority of PacketGame in end-to-end concurrency. Mu Yuan, Lan Zhang 0002, Xuanke You, Xiang-Yang Li 0001 |
SIGCOMM | 1 |
| 2023 | CoTel: Ontology-Neural Co-Enhanced Text LabelingabstractThe success of many web services relies on the large-scale domain-specific high-quality labeled dataset. Insufficient public datasets motivate us to reduce the cost of data labeling while maintaining high accuracy in support of intelligent web applications. The rule-based method and the learning-based method are common techniques for labeling. In this work, we study how to utilize the rule-based and learning-based methods for resource-effective text labeling. We propose CoTel, the first ontology-neural co-enhanced framework for text labeling. We propose critical ontology extraction in the rule-based module and ontology-enhanced loss prediction in the learning-based module. CoTel can integrate explicit labeling rules and implicit labeling models and make them help each other to improve resource efficiency in text labeling tasks. We evaluate CoTel on both public datasets and real applications with three different tasks. Compared with the baseline, CoTel can reduce the time cost by 64.75% (a 2.84× speedup) and the number of labeling by 62.07%. Miaohui Song, Lan Zhang 0002, Mu Yuan, Zichong Li, Qi Song 0004, Yijun Liu 0003, Guidong Zheng |
WWW | 3 |
| 2023 | MLink: Linking Black-Box Models From Multiple Domains for Collaborative InferenceabstractThe cost efficiency of model inference is critical to real-world machine learning (ML) applications, especially for delay-sensitive tasks and resource-limited devices. A typical dilemma is: in order to provide complex intelligent services (e.g., smart city), we need inference results of multiple ML models, but the cost budget (e.g., GPU memory) is not enough to run all of them. In this work, we study underlying relationships among black-box ML models and propose a novel learning task: model linking, which aims to bridge the knowledge of different black-box models by learning mappings (dubbed model links) between their output spaces. We propose the design of model links which supports linking heterogeneous black-box ML models. Also, in order to address the distribution discrepancy challenge, we present adaptation and aggregation methods of model links. Based on our proposed model links, we developed a scheduling algorithm, named MLink. Through collaborative multi-model inference enabled by model links, MLink can improve the accuracy of obtained inference results under the cost budget. We evaluated MLink on a multi-modal dataset with seven different ML models and two real-world video analytics systems with six ML models and 3,264 hours of video. Experimental results show that our proposed model links can be effectively built among various black-box models. Under the budget of GPU memory, MLink can save 66.7% inference computations while preserving 94% inference accuracy, which outperforms multi-task learning, deep reinforcement learning-based scheduler and frame filtering baselines. Mu Yuan, Lan Zhang 0002, Zimu Zheng, Xiang-Yang Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Several Families of Binary Minimal Linear Codes From Two-to-One FunctionsabstractMinimal linear codes have important applications in secure communications, including in the framework of secret sharing schemes and secure multi-party computation. A lot of research have been carried out to derive codes with few weights (but more importantly, being minimal) using algebraic or geometric approaches. One of the main power and fructify algebraic methods is based on the design of those codes by employing functions over finite fields. Li et al. (2021) have recently identified some binary linear codes with few weights from two classes of two-to-one functions. In this paper, our ultimate objective is to expand the class of codes derived from the paper of Li et al. by proposing larger classes of binary linear codes with few weights via generic constructions involving other known families of two-to-one functions over the finite field$\mathbb {F}_{2^{n}}$of order$2^{n}$. We succeed in constructing such codes, and we also completely determine their weight distributions. The linear codes presented in this paper differ in parameters from those known in the literature. Besides, some of them are optimal concerning the well-known Griesmer bound. Notably, we prove that our codes are either optimal or almost optimal with respect to the online Database of Grassl. We next observe that the derived binary linear codes also have the minimality property for most cases. We then describe the access structures of the secret-sharing schemes based on their dual codes. Finally, we solve two problems left open in the paper by Li et al. (more specifically, a complete solution to Problem 2 and a partial solution to Problem 1). Sihem Mesnager, Liqin Qian, Xiwang Cao, Mu Yuan |
IEEE Trans. Inf. Theory | 4 |
| 2023 | More About the Corpus of Involutions From Two-to-One Mappings and Related Cryptographic S-BoxesabstractPermutation polynomials have been extensively studied for their applications in cryptography, coding theory, combinatorial design, etc. An important subfamily of permutations is the class of involutions (those permutations are equal to their compositional inverse). Elements of this class have been used frequently for block cipher designs and coding theory. In this article, we further investigate this corpus using new approaches, specifically from two-to-one (2-to-1) functions and (in some cases) using the graph indicators introduced by Carlet in 2020. In our constructions of involutions over the finite field$\mathbb {F}_{2^{n}}$of order$2^{n}$, we shall intensively use 2-to-1 mappings over$\mathbb {F}_{2^{n}}$. More specifically, we present a new constructive method to design involutions from 2-to-1 mappings through their graph indicator and derive new involutions from known 2-to-1 mappings. Besides, we also propose several new classes of 2-to-1 mappings, including 2-to-1 hexanomials, 2-to-1 mappings of the form$(x^{2^{k}}+x+\delta)^{s_{1}}+(x^{2^{k}}+x+\delta)^{s_{2}}+cx$, and 2-to-1 mappings from linear 2-to-1 mappings. We also exhibit the corresponding involutions of the constructed 2-to-1 mappings. Furthermore, an infinite family of involutions with differential uniformity at most 4 (EA-inequivalent to the inverse function) is obtained. Finally, we highlight that all our derived families of involutions have no fixed point, further accentuating their cryptographic interest. Sihem Mesnager, Mu Yuan, Dabin Zheng |
IEEE Trans. Inf. Theory | 2 |
| 2023 | MultiSense: Cross-labelling and Learning Human Activities Using Multimodal Sensing DataabstractTo tap into the gold mine of data generated by Internet of Things (IoT) devices with unprecedented volume and value, there is an urgent need to efficiently and accurately label raw sensor data. To this end, we explore and leverage the hidden connections among the multimodal data collected by various sensing devices and propose to let different modal data complement and learn from each other. But it is challenging to align and fuse multimodal data without knowing their perception (and thus the correct labels). In this work, we propose MultiSense , a paradigm for automatically mining potential perception, cross-labelling each modal data, and then updating the learning models for recognizing human activity to achieve higher accuracy or even recognize new activities. We design innovative solutions for segmenting, aligning, and fusing multimodal data from different sensors, as well as model updating mechanism. We implement our framework and conduct comprehensive evaluations on a rich set of data. Our results demonstrate that MultiSense significantly improves the data usability and the power of the learning models. With nine diverse activities performed by users, our framework automatically labels multimodal sensing data generated by five different sensing mechanisms (video, smart watch, smartphone, audio, and wireless-channel) with an average accuracy 98.5%. Furthermore, it enables models of some modalities to learn unknown activities from other modalities and greatly improves the activity recognition ability. Lan Zhang 0002, Daren Zheng, Mu Yuan, Zhengtao Wu, Mengjing Liu, Xiang-Yang Li 0001 |
ACM Trans. Sens. Networks | 3 |
| 2022 | MLink: Linking Black-Box Models for Collaborative Multi-Model InferenceabstractThe cost efficiency of model inference is critical to real-world machine learning (ML) applications, especially for delay-sensitive tasks and resource-limited devices. A typical dilemma is: in order to provide complex intelligent services (e.g. smart city), we need inference results of multiple ML models, but the cost budget (e.g. GPU memory) is not enough to run all of them. In this work, we study underlying relationships among black-box ML models and propose a novel learning task: model linking. Model linking aims to bridge the knowledge of different black-box models by learning mappings (dubbed model links) between their output spaces. Based on model links, we developed a scheduling algorithm, named MLink. Through collaborative multi-model inference enabled by model links, MLink can improve the accuracy of obtained inference results under the cost budget. We evaluated MLink on a multi-modal dataset with seven different ML models and two real-world video analytics systems with six ML models and 3,264 hours of video. Experimental results show that our proposed model links can be effectively built among various black-box models. Under the budget of GPU memory, MLink can save 66.7% inference computations while preserving 94% inference accuracy, which outperforms multi-task learning, deep reinforcement learning-based scheduler and frame filtering baselines. Mu Yuan, Lan Zhang 0002, Xiang-Yang Li 0001 |
AAAI | 1 |
| 2022 | InFi: end-to-end learnable input filter for resource-efficient mobile-centric inferenceabstractMobile-centric AI applications put forward high requirements for resource-efficiency of model inference. Input filtering is a promising approach to eliminate the redundancy in the input so as to reduce the cost of inference. Previous efforts have tailored effective solutions for many applications, but left two essential questions unanswered: (1) theoretical filterability of an inference workload to guide the application of input filtering techniques, thereby avoiding the trial-and-error cost for resource-constrained mobile applications; (2) robust discriminability of feature embedding to allow input filtering to be widely effective for diverse inference tasks and input content. To answer these questions, we first provide a generic formalization of the input filtering problem and theoretically compare the hypothesis complexity of inference models and their input filters to understand the optimization potential of applying input filtering. Then we propose the first end-to-end learnable input filtering framework that covers most state-of-the-art methods and surpasses them in feature embedding with robust discriminability. Based on our framework, we design and implement an input filtering system InFi supporting six input modalities. InFi is the first to support text and sensor signal inputs and model partitioning deployments widely adopted by under-resourced mobile systems. Comprehensive evaluations confirm our theoretical results and show that InFi outperforms strong baselines in applicability, accuracy, and efficiency, owing to its generality and end-to-end learnability. InFi can achieve 8.5X throughput and save 95% bandwidth, while keeping over 90% accuracy, for a video analytics app on mobile platforms. Mu Yuan, Lan Zhang 0002, Fengxiang He, Xueting Tong, Xiang-Yang Li 0001 |
MobiCom | 1 |
| 2022 | Adaptive Model Scheduling for Resource-efficient Data LabelingabstractLabeling data (e.g., labeling the people, objects, actions, and scene in images) comprehensively and efficiently is a widely needed but challenging task. Numerous models were proposed to label various data and many approaches were designed to enhance the ability of deep learning models or accelerate them. Unfortunately, a single machine-learning model is not powerful enough to extract various semantic information from data. Given certain applications, such as image retrieval platforms and photo album management apps, it is often required to execute a collection of models to obtain sufficient labels. With limited computing resources and stringent delay, given a data stream and a collection of applicable resource-hungry deep-learning models, we design a novel approach to adaptively schedule a subset of these models to execute on each data item, aiming to maximize the value of the model output (e.g., the number of high-confidence labels). Achieving this lofty goal is nontrivial since a model’s output on any data item is content-dependent and unknown until we execute it. To tackle this, we propose an Adaptive Model Scheduling framework, consisting of (1) a deep reinforcement learning-based approach to predict the value of unexecuted models by mining semantic relationship among diverse models, and (2) two heuristic algorithms to adaptively schedule the model execution order under a deadline or deadline-memory constraints, respectively. The proposed framework does not require any prior knowledge of the data, which works as a powerful complement to existing model optimization technologies. We conduct extensive evaluations on five diverse image datasets and 30 popular image labeling models to demonstrate the effectiveness of our design: our design could save around 53% execution time without loss of any valuable labels. Mu Yuan, Lan Zhang 0002, Xiang-Yang Li 0001, Linzhuo Yang, Hui Xiong 0001 |
ACM Trans. Knowl. Discov. Data | 1 |
| 2021 | M&M: Recognizing Multiple Co-evolving Activities From Multi-source VideosabstractThe wide deployment of surveillance systems has shown the necessity to recognize human activities in videos. Existing work achieved high recognition accuracy for single- person activities and some group activities. In this paper, we identify a more challenging issue: recognizing multiple co- evolving but asynchronous activities from a set of videos captured by multiple cameras that share overlapped views. To address this issue, we design a system M&M to fuse objects and 2D skeletons from multi-source videos to reconstruct the complete 3D view-invariant model of multiple activity scenarios, which enables accurate recognition of cross-source activities. By embedding the 3D model with a concise graph representation, we propose an efficient recognition method in a bottom-up manner to achieve high accuracy and good scalability to changing complexity of the captured scenario. We collect and release a dataset containing correlated multi-source videos for multiple co-evolving activities, and evaluate our design on it. Experimental results show that M&M achieves 91.2% average accuracy for 10 types of coevolving activities. Lan Zhang 0002, Mu Yuan, Daren Zheng, Xiang-Yang Li 0001 |
DCOSS | 2 |
| 2021 | MultiSense: Cross Labelling and Learning Human Activities Using Multimodal Sensing DataabstractOne of the major challenges for fully enjoying the power of machine learning is the need for the high-quality labelled data. To tap-in the gold-mine of data generated by IoT devices with unprecedented volume and value, we discover and leverage the hidden connections among the multimodal data collected by various sensing devices. Different modal data can complement and learn from each other, but it is challenging to fuse multimodal data without knowing their perception (and thus the correct labels). In this work, we propose MultiSense, a paradigm for automatically mining potential perception, cross labelling each modal data, and then improving the learning models for recognizing human activity accurately. We design innovative solutions for segmenting, aligning, and fusing multimodal data from different sensors. We implement our framework and conduct comprehensive evaluations on a rich set of data. Our results demonstrate that MultiSense significantly improves the data usability and the power of the learning models. With 9 diverse activities performed by users, our framework automatically labels multimodal sensing data generated by five different sensing mechanisms (video, smart watch, smartphone, audio, and wireless-channel) with an average accuracy 98.5%, while for each single-modal data model, the accuracy is 95.2%, 92.1%, 71.1%, 91.5%, and 34.8% respectively. Furthermore, it enables the existing models to learn unknown activities from other modalities and thus greatly improves the activity recognition accuracy. Lan Zhang 0002, Daren Zheng, Zhengtao Wu, Mengjing Liu, Mu Yuan, Xiang-Yang Li 0001 |
MASS | 5 |
| 2020 | Comprehensive and Efficient Data Labeling via Adaptive Model SchedulingabstractLabeling data comprehensively and efficiently is a widely needed but challenging task. With limited computing resources, given a data stream and a collection of deep-learning models, we propose to adaptively select and schedule a subset of these models to execute, aiming to maximize the value of the model output. Achieving this goal is nontrivial since a model's output on any data item is content-dependent and hard to predict. In this paper, we present an Adaptive Model Scheduling framework, consisting of 1) a deep reinforcement learning-based approach to predict the value of unexecuted models by mining semantic relationship among diverse models, and 2) two heuristic algorithms to adaptively schedule models under deadline or deadline-memory constraints. The proposed framework does not require any prior knowledge of the data, which works as a powerful complement to existing model optimization technologies. We conduct extensive evaluations on 30 popular image labeling models to demonstrate the effectiveness of our design. Mu Yuan, Lan Zhang 0002, Xiang-Yang Li 0001, Hui Xiong 0001 |
ICDE | 1 |
| 2020 | High-quality Activity-Level Video AdvertisingabstractOnline video advertising is a billion dollar business, but the current low CTR reveals the huge potential for improvement in the ad serving quality. In this work, we present a novel activity-level video advertising system named ActVa. Different from existing systems that assume a fixed scope of ad keywords, ActVA enables advertising targeted to non-predefined activities in a highly efficient way requiring no training data for diverse activities. To achieve this goal, a general and extensible graphical representation of both video content and advertising demand is proposed to embed multimodal content at the activity level. Our ad-content relevancy measurement can achieve 10,000 FPS retrieval speed. We model the ads assigning task as an optimization problem taking content relevance, ads revenue as well as viewer experience into consideration. A non-maximal suppression based algorithm is designed to significantly reduce the algorithm complexity for online ad serving. Our extensive objective and subjective experimental results show the effectiveness and efficiency of ActVA. ActVA can effectively uncovers numerous high-quality (content-relevant) advertising opportunities and delivers ads to viewers in a profitable and user-friendly way. Mu Yuan, Lan Zhang 0002, Zhengtao Wu, Daren Zheng |
IWQoS | 1 |
| 2019 | Poster: Cross Labelling and Learning Unknown Activities Among Multimodal Sensing DataabstractOne of the major challenges for fully enjoying the power of machine learning is the need for the high-quality labelled data. To tap-in the gold-mine of data generated by IoT devices with unprecedented volume and value, we discover and leverage the hidden connections among the multimodal data collected by various sensing devices. Different modal data can complete and learn from each other, but it is challenging to fuse multimodal data without knowing their perception (and thus the correct labels). In this work, we propose MultiSense, a paradigm for automatically mining potential perception, cross-labelling each modal data, and then improving the learning models over the set of multimodal data. We design innovative solutions for segmenting, aligning, and fusing multimodal data from different sensors. We implement our framework and conduct comprehensive evaluations on a rich set of data. Our results demonstrate that MultiSense significantly improves the data usability and the power of the learning models. Lan Zhang 0002, Daren Zheng, Zhengtao Wu, Mengjing Liu, Mu Yuan, Xiang-Yang Li 0001 |
MobiCom | 5 |
| 2019 | Constructions of Involutions Over Finite FieldsabstractAn involution over finite fields is a permutation polynomial whose inverse is itself. Owing to this property, involutions over finite fields have been widely used in applications, such as cryptography and coding theory. Following the idea by Wang to characterize the involutory behavior of the generalized cyclotomic mappings, this paper gives a more concise criterion for$x^{r}h(x^{s})\in {\mathbb F} _{q}[x]$being involutions over the finite field${\mathbb F}_{q}$, where$r\geq 1$and$s\,|\, (q-1)$. By using this criterion, we propose a general method to construct involutions of the form$x^{r}h(x^{s})$over${\mathbb F}_{q}$from given involutions over some subgroups of${\mathbb F}_{q}^{*}$by solving congruent and linear equations over finite fields. Then, many classes of explicit involutions of the form$x^{r}h(x^{s})$over${\mathbb F}_{q}$are obtained. Dabin Zheng, Mu Yuan, Nian Li 0005, Lei Hu 0003, Xiangyong Zeng |
IEEE Trans. Inf. Theory | 2 |
| 2010 | User-preference-based service selection using fuzzy logicabstractWeb services provide a standard interaction interface for network-based services. With so many diverse web services available over the Internet, services can be composed together and delivered to users as a package. The way to combine services is still an open problem, especially considering users' preferences. Users always have vague opinions on these preferences when they choose component services. In this paper, we propose a user-preference-based selection engine to compose services, which allows users to define non-quantifiable factors and policies to represent their preferences. Then, the engine automatically composes web services into a package following these policies by use fuzzy logic. The engine ensures the whole package has the most satisfaction. Zhengping Wu, Mu Yuan |
CNSM | 2 |