Chaoyue Niu

dblp:176/5779 · DBLP profile ↗
← Back
39ranked-venue papers
13as first author
30since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 2 first-author · 14 since 2021Databases, data management, data science and information retrieval · 15 · 6 first-author · 11 since 2021Computer networks · 7 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Systems, architecture and hardware · 4 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Automated Annotation of Privacy Information in User Interactions with Large Language Models
Chaoyue Niu, Fan Wu 0006, Shaojie Tang 0001, Guihai Chen
KDD (1)4
2026 Collaborative Service of On-Device Unimodal Model and Cloud-Based Multimodal Model for Mobile Livestreaming Content Understanding
abstract
To guide consumers to mobile livestreams that involve their interested products, it is necessary to have a good understanding of livestreaming contents, and the key task is to accurately recognize the products being promoted by streamers. The mainstream cloud-based service framework is challenged by the high concurrency of requests, the high overhead of multimodal recognition, and the requirement of low latency. To break these bottlenecks, we propose a new device-cloud collaborative service framework, where each streamer's mobile device holds unimodal model that can process most of video frames and also uploads the extracted unimodal features to facilitate the cloud-side multimodal recognition of the remaining few frames. On-device unimodal model is further incrementally trained over the samples constructed by leveraging the streamers' manual labeling behaviors, leading to personalized version that can adapt to the heterogeneous and dynamic livestreaming contents of different streamers. Nevertheless, the device-side personalized unimodal features are misaligned in feature space and cannot be directly fused into the cloud-side multimodal model. We thus design a pluggable prompt generation module to transform the personalized unimodal features into prompt embeddings and prepend them to the original inputs of the multimodal backbone network, instructing feature fusion in a streamer-specific manner. We extensively evaluate on the public Fashion-Gen dataset and an industrial dataset collected from Taobao Live. We also perform online testing and practical overhead testing in Taobao Live. Evaluation results reveal the effectiveness and efficiency of our design as well as the consistent advantage over existing baselines.
Chaoyue Niu, Yutong Dai 0005, Yikai Yan, Zhijie Cao, Chengfei Lv, Shaojie Tang 0001, Fan Wu 0006, Guihai Chen
IEEE Trans. Serv. Comput.1
2025 Querier-Aware LLM: Generating Personalized Responses to the Same Query from Different Queriers
abstract
Existing work on large language model (LLM) personalization assigned different responding roles to LLMs, but overlooked the diversity of queriers. In this work, we propose a new form of querier-aware LLM personalization, generating different responses even for the same query from different queriers. We design a dual-tower model architecture with a cross-querier general encoder and a querier-specific encoder. We further apply contrastive learning with multi-view augmentation, pulling close the dialogue representations of the same querier, while pulling apart those of different queriers. To mitigate the impact of query diversity on querier-contrastive learning, we cluster the dialogues based on query similarity and restrict the scope of contrastive learning within each cluster. To address the lack of datasets designed for querier-aware personalization, we also build a multi-querier dataset from English and Chinese scripts, as well as WeChat records, called MQDialog, containing 173 queriers and 12 responders. Extensive evaluations demonstrate that our design significantly improves the quality of personalized response generation, achieving relative improvement of 8.4% to 48.7% in ROUGE-L scores and winning rates ranging from 54% to 82% compared with various baseline methods.
Chaoyue Niu, Fan Wu 0006, Chengfei Lv, Guihai Chen
CIKM2
2025 Adaptive Routing of Text-to-Image Generation Requests between Large Cloud Model and Light-Weight Edge Model
Zewei Xin, Qinya Li, Chaoyue Niu, Fan Wu 0006, Guihai Chen
ICCV3
2025 Personalized Language Model Learning on Text Data Without User Identifiers
abstract
In many practical natural language applications, user data are highly sensitive, requiring anonymous uploads of text data from mobile devices to the cloud without user identifiers. However, the absence of user identifiers restricts the ability of cloud-based language models to provide personalized services, which are essential for catering to diverse user needs. The trivial method of replacing an explicit user identifier with a static user embedding as model input still compromises data anonymization. In this work, we propose to let each mobile device maintain a user-specific distribution to dynamically generate user embeddings, thereby breaking the one-to-one mapping between an embedding and a specific user. We further theoretically demonstrate that to prevent the cloud from tracking users via uploaded embeddings, the local distributions of different users should either be derived from a linearly dependent space to avoid identifiability or be close to each other to prevent accurate attribution. Evaluation on both public and industrial datasets using different language models reveals a remarkable improvement in accuracy from incorporating anonymous user embeddings, while preserving real-time inference requirement.
Yangwenjian Tan, Chaoyue Niu, Fandong Meng, Jie Zhou 0016, Fan Wu 0006, Guihai Chen
KDD (1)4
2025 Real-Time Drone Flight Height Decision for 3D Building Reconstruction with NeRF
abstract
Drones serve as the primary camera-mounted platform to collect high-resolution images of buildings from various viewpoints for 3D neural radiance fields (NeRF) reconstruction. However, the flexibility of drones in lowering flight heights to supplement richer building details and thus improve the reconstruction quality of NeRF has not been explored. Given the limited power of a drone and the rapidly changing nature of outdoor scenes, it is necessary to quickly predict an optimal supplementary lower flight height after the drone finishes capturing at a default height. The main challenges involve the implicit relationship between supplementary flight heights and reconstruction quality, the complexity of prediction task with little prior information, and the tension between the need for real-time decision making and the resource constraints inherent to on-drone task execution. In this work, we develop an end-to-end system pipeline with an offline-online decoupling feature. We first design a model architecture that embeds supplementary flight heights to predict the reconstruction quality of NeRF using images captured at the default height. To enhance model generalization ability in a cost-effective manner, we offline train an ensemble of multiple models over the samples constructed from sandbox buildings using building-level cross validation. During the online serving phase on the drone, the most appropriate model is selected based on feature matching between real and sandbox buildings to determine the optimal supplementary flight height. We build a testbed and evaluate on 16 sandbox buildings and 8 real-world buildings. Evaluation results demonstrate that our pipeline accurately identifies the optimal supplementary lower height in a few minutes and improves reconstruction quality.
Shenghao Jia, Chaoyue Niu, Fan Wu 0006
MASS2
2025 CEFSW'25: The 2nd Collaboration and Evolution of Foundation and Specialized Models Workshop
abstract
Foundation models (FMs), known for their broad cognitive capabilities but often constrained to cloud deployment, and specialized models (SMs), characterized by their lightweight, goal-oriented nature suitable for devices, offer complementary strengths. Traditional cloud-centric paradigms face limitations in real-time performance, personalization, cost, and privacy, highlighting the need for innovative approaches that leverage device-level capabilities. This workshop served as a platform to discuss the rapid advancements and emerging research directions in FM-SM collaboration and co-evolution. Key focus areas included: (i) novel collaborative frameworks bridging cloud FMs and device SMs, (ii) mechanisms for model evolution, knowledge transfer, aggregation, and generation, (iii) integration of multimodal perspectives, particularly for multimedia retrieval tasks relevant to ICMR, (iv) strategies for enhancing robustness, interpretability, and fairness, and (v) the development of new benchmarks and resources. Featuring keynote presentations and peer-reviewed papers on topics ranging from multimodal understanding and reasoning to efficient on-device fine-tuning and mobile agents, the workshop fostered interdisciplinary dialogue.
Shengyu Zhang 0001, Fan Yao 0002, Chaoyue Niu, Hongxia Yang, Fan Wu 0006, Fei Wu 0001
ICMR4
2025 CORE: Reducing UI Exposure in Mobile Agents via Collaboration Between Cloud and Local LLMs
abstract
Mobile agents rely on Large Language Models (LLMs) to plan and execute tasks on smartphone user interfaces (UIs). While cloud-based LLMs achieve high task accuracy, they require uploading the full UI state at every step, exposing unnecessary and often irrelevant information. In contrast, local LLMs avoid UI uploads but suffer from limited capacity, resulting in lower task success rates. We propose $\textbf{CORE}$, a $\textbf{CO}$llaborative framework that combines the strengths of cloud and local LLMs to $\textbf{R}$educe UI $\textbf{E}$xposure, while maintaining task accuracy for mobile agents. CORE comprises three key components: (1) $\textbf{Layout-aware block partitioning}$, which groups semantically related UI elements based on the XML screen hierarchy; (2) $\textbf{Co-planning}$, where local and cloud LLMs collaboratively identify the current sub-task; and (3) $\textbf{Co-decision-making}$, where the local LLM ranks relevant UI blocks, and the cloud LLM selects specific UI elements within the top-ranked block. CORE further introduces a multi-round accumulation mechanism to mitigate local misjudgment or limited context. Experiments across diverse mobile apps and tasks show that CORE reduces UI exposure by up to 55.6\% while maintaining task success rates slightly below cloud-only agents, effectively mitigating unnecessary privacy exposure to the cloud. The code is available at https://github.com/Entropy-Fighter/CORE.
Gucongcong Fan, Chaoyue Niu, Chengfei Lyu, Fan Wu 0006, Guihai Chen
NeurIPS2
2025 RAGRouter: Learning to Route Queries to Multiple Retrieval-Augmented Language Models
abstract
Retrieval-Augmented Generation (RAG) significantly improves the performance of Large Language Models (LLMs) on knowledge-intensive tasks. However, varying response quality across LLMs under RAG necessitates intelligent routing mechanisms, which select the most suitable model for each query from multiple retrieval-augmented LLMs via a dedicated router model. We observe that external documents dynamically affect LLMs' ability to answer queries, while existing routing methods, which rely on static parametric knowledge representations, exhibit suboptimal performance in RAG scenarios. To address this, we formally define the new retrieval-augmented LLM routing problem, incorporating the influence of retrieved documents into the routing framework. We propose RAGRouter, a RAG-aware routing design, which leverages document embeddings and RAG capability embeddings with contrastive learning to capture knowledge representation shifts and enable informed routing decisions. Extensive experiments on diverse knowledge-intensive tasks and retrieval settings, covering open and closed-source LLMs, show that RAGRouter outperforms the best individual LLM and existing routing methods. With an extended score-threshold-based mechanism, it also achieves strong performance-efficiency trade-offs under low-latency constraints. The code and data are available at https://github.com/OwwO99/RAGRouter.
Chaoyue Niu, Fan Wu 0006, Guihai Chen
NeurIPS4
2025 HRM^2Avatar: High-Fidelity Real-Time Mobile Avatars from Monocular Phone Scans
abstract
We present HRM2Avatar, a novel framework for creating high-fidelity avatars from monocular phone scans, which can be rendered and animated in real-time on mobile devices. Monocular capture with commodity smartphones provides a low-cost, pervasive alternative to studio-grade multi-camera rigs, making avatar digitization accessible to non-expert users. Reconstructing high-fidelity avatars from single-view video sequences poses significant challenges due to deficient visual and geometric data relative to multi-camera setups. To address these limitations, at the data level, our method leverages two types of data captured with smartphones: static pose sequences for detailed texture reconstruction and dynamic motion sequences for learning pose-dependent deformations and lighting changes. At the representation level, we employ a lightweight yet expressive representation to reconstruct high-fidelity digital humans from sparse monocular data. First, we extract explicit garment meshes from monocular data to model clothing deformations more effectively. Second, we attach illumination-aware Gaussians to the mesh surface, enabling high-fidelity rendering and capturing pose-dependent lighting changes. This representation efficiently learns high-resolution and dynamic information from our tailored monocular data, enabling the creation of detailed avatars. At the rendering level, real-time performance is critical for rendering and animating high-fidelity avatars in AR/VR, social gaming, and on-device creation, demanding sub-frame responsiveness. Our fully GPU-driven rendering pipeline delivers 120 FPS on mobile devices and 90 FPS on standalone VR devices at 2K resolution, over 2.7 × faster than representative mobile-engine baselines. Experiments show that HRM2Avatar delivers superior visual realism and real-time interactivity at high resolutions, outperforming state-of-the-art monocular methods.
Shenghao Jia, Liangchao Zhu, Zhonglei Yang, Jinze Ma, Chaoyue Niu, Chengfei Lv
SIGGRAPH Asia8
2025 Federated multi-task learning with cross-device heterogeneous task subsets
Zewei Xin, Qinya Li, Chaoyue Niu, Fan Wu 0006, Guihai Chen
J. Parallel Distributed Comput.3
2025 ARSys: An Efficient and Cross-Platform Development, Deployment, and Runtime System for Mobile Augmented Reality
abstract
Augmented reality (AR) offers users immersive experiences to interact with digital contents in their physical space. However, practical AR applications are challenged by the tight coupling of algorithm and engineering during the development and deployment phases as well as the execution requirements of hybrid AR subtasks on heterogeneous and resource-constraint mobile devices. In this work, we build an end-to-end, cross-platform, and efficient AR system, called ARSys. The infrastructure in ARSys adopts the new principle of integrated design, unifies and refines AR fundamental capabilities, supports streaming media processing, model inference, and real-time rendering by exposing high-performance tensor compute engine to top, and constructs a Python multi-instance virtual machine as the cross-platform AR task execution container. The runtime mechanism of ARSys schedules AR tasks in a pipeline parallelism way and allocates subtasks to hardware backends by optimizing the slowest node. The development workbench and the deployment platform in ARSys allow the decoupling of algorithms written in Python from engineering components in C/C++ and further support remote debugging and quick validation of AR algorithms. We extensively evaluate ARSys in practical AR applications across high-end, mid-end, and low-end Android and iOS devices, demonstrating higher development, deployment, and runtime efficiency than existing MediaPipe-oriented framework. ARSys has been integrated into Mobile Taobao for production use.
Chengfei Lv, Chaoyue Niu, Xiaotang Jiang, Fan Wu 0006, Guihai Chen
IEEE Trans. Mob. Comput.2
2024 Adapting the Attention of Cloud-Based Recognition Model to Client-Side Images without Local Re-Training
Yangwenjian Tan, Yikai Yan, Chaoyue Niu
ACML3
2024 MPOD123: One Image to 3D Content Generation Using Mask-Enhanced Progressive Outline-to-Detail Optimization
abstract
Recent advancements in single image driven 3D content generation have been propelled by leveraging prior knowledge from pretrained 2D diffusion models. However, the 3D content generated by existing methods often exhibits distorted outline shapes and inadequate details. To solve this problem, we propose a novel framework called Mask-enhanced Progressive Outline-to-Detail optimization (aka. MPOD123), which consists of two stages. Specifically, in the first stage, MPOD123 utilizes the pretrained view-conditioned diffusion model to guide the outline shape optimization of the 3D content. Given certain viewpoint, we estimate outline shape priors in the form of 2D mask from the 3D content by leveraging opacity calculation. In the second stage, MPOD123 incorporates Detail Appearance Inpainting (DAI) to guide the refinement on local geometry and texture with the shape priors. The essence of DAI lies in the Mask Rectified Cross-Attention (MRCA), which can be conveniently plugged in the stable diffusion model. The MRCA module utilizes the mask to rectify the attention map from each cross-attention layer. Accompanied with this new module, DAI is capable of guiding the detail refinement of the 3D content, while better preserves the outline shape. To assess the applicability in practical scenarios, we contribute a new dataset modeled on real-world e-commerce environments. Extensive quantitative and qualitative experiments on this dataset and open benchmarks demonstrate the effectiveness of MPOD123 over the state-of-the-arts.
Jimin Xu, Tianbao Wang, Tao Jin 0004, Shengyu Zhang 0001, Jiangjing Lyu, Chengfei Lv, Chaoyue Niu, Zhou Yu 0001, Zhou Zhao 0001, Fei Wu 0001
CVPR9
2024 Picking Models for Heterogeneous Clients: A Server-Client Feature Contrastive Learning Design
abstract
Conventional deep learning applications for clients generally adapt a global model trained on large-scale server data to different clients, or rely on meta-learning methods to be client-adaptive. However, cross-client data heterogeneity and the building of meta-training resources on server result in degraded client performance and excessive server overhead. In this work, we propose ContrastPick, a design friendly to both sides of the server-client scenario, to provide proper deep learning models to different clients. Given a large-scale dataset and a collection of candidate models from a pretrained supernet on server, we separate the server dataset into smaller subsets and then evaluate the candidates on such subsets. We introduce a contrastive approach to learn a dataset encoder to minimize the similarity between separated subset pairs, thereby matching a client dataset to a suitable subset with high-performing candidate models. With ContrastPick, we can efficiently and effectively serve models to clients without introducing massive re-training and storage overhead caused by meta-training resources. Extensive experimental results on 4 datasets composed of heterogeneous clients demonstrate that ContrastPick can achieve superior performance with significantly reduced overhead compared with baselines.
Chaoyue Niu, Fan Wu 0006
HPCC2
2024 Pixel-based Hole Quality Evaluation in Robot Drilling Manufacturing Process
abstract
Aircraft assembly entails drilling numerous holes, often in multi-material stacks, then joining parts with fasteners fitted through the holes. There are stringent quality requirements on the holes, and assessment of hole quality is crucial to ensuring the integrity of the joints. Carbon fibre reinforced polymer (CFRP) is a commonly used material in aircraft structures due to its desirable properties. However, it is susceptible to defects not associated with metals, such as delamination and uncut fibres. While there have been multiple metrics for assessment of these defects proposed in literature, relatively few attempts to consolidate them have been seen. Furthermore, common measurement methods used for assessing these defects (e.g. 3D-microscopy) are well established, but can be time-consuming; manual interrogation of raw inspection data can also be highly subjective. To address these challenges, this paper proposes a set of combined metrics along with an automatic image processing framework for fast assessment of delamination and uncut fibre defects. The combined metric for the delamination region is aggregated from across six factors and uncut fibre uniform metrics from across five factors. The image processing framework receives grayscale images as inputs, taken from an optical coordinate measuring machine, and outputs the combined metrics along with eleven separate metrics. Experimental results, using a preexisting dataset from twenty-four holes on two workpieces from a real robotic drilling operation, are given to demonstrate the effectiveness of the proposed combined metrics and the corresponding image processing framework.
Chaoyue Niu, Erica Smith, Robert Bramley, Pete Crawforth, Mahdi Mahfouf, Visakan Kadirkamanathan
INDIN1
2024 Enhancing On-Device LLM Inference with Historical Cloud-Based LLM Interactions
abstract
Many billion-scale large language models (LLMs) have been released for resource-constraint mobile devices to provide local LLM inference service when cloud-based powerful LLMs are not available. However, the capabilities of current on-device LLMs still lag behind those of cloud-based LLMs, and how to effectively and efficiently enhance on-device LLM inference becomes a practical requirement. We thus propose to collect the user's historical interactions with the cloud-based LLM and build an external datastore on the mobile device for enhancement using nearest neighbors search. Nevertheless, the full datastore improves the quality of token generation at the unacceptable expense of much slower generation speed. To balance performance and efficiency, we propose to select an optimal subset of the full datastore within the given size limit, the optimization objective of which is proven to be submodular. We further design an offline algorithm, which selects the subset after the construction of the full datastore, as well as an online algorithm, which performs selection over the stream and can be flexibly scheduled. We theoretically analyze the performance guarantee and the time complexity of the offline and the online designs to demonstrate effectiveness and scalability. We finally take three ChatGPT related dialogue datasets and four different on-device LLMs for evaluation. Evaluation results show that the proposed designs significantly enhance LLM performance in terms of perplexity while maintaining fast token generation speed. Practical overhead testing on the smartphone reveal the efficiency of on-device datastore subset selection from memory usage and computation overhead.
Chaoyue Niu, Fan Wu 0006, Shaojie Tang 0001, Chengfei Lyu, Guihai Chen
KDD2
2024 An End-to-End, Low-Cost, and High-Fidelity 3D Video Pipeline for Mobile Devices
abstract
To provide full-body 3D videos of performers showcasing diverse clothes with dynamic movements for e-commerce platforms, we develop an end-to-end, low-cost, and high-fidelity production and deployment pipeline. We first set up a low-cost capture studio with only 24 RGB cameras and embrace fast neural surface reconstruction to produce high-quality meshes without depth information. We then quickly group all the frames with local motion priors, select a keyframe for each group, and accurately register any other frame to the keyframe under the guidance of semantic labels, thereby avoiding transmitting all the frames to mobile devices and loading them into memory. For real-time rendering, we propose an on-device sparse computation method for efficient deformation from keyframes to the other frames. Evaluation over 2 self-captured performances and 8 public performances reveals that the pipeline achieves the reconstruction time of 28 minutes per frame, the average PSNR of 30.4, the average bandwidth requirement of 4.2MB/s, and the on-device frame rate of 60 fps, demonstrating superiority over existing baselines.
Tiancheng Fang, Chaoyue Niu, Yujie Sun 0001, Chengfei Lv, Xiaotang Jiang, Ben Xue, Fan Wu 0006, Guihai Chen
MobiCom2
2024 Federated Optimization Under Intermittent Client Availability
abstract
Federated learning is a new distributed machine learning framework, where numerous heterogeneous clients collaboratively train a model without sharing training data. In this work, we consider a practical and ubiquitous issue when deploying federated learning in mobile environments: intermittent client availability, where the set of eligible clients may change during the training process. Such intermittent client availability would seriously deteriorate the performance of the classical federated averaging algorithm (FedAvg). Thus, we propose a simple distributed nonconvex optimization algorithm, called federated latest averaging (FedLaAvg), which leverages the latest gradients of all clients, even when the clients are not available, to jointly update the global model in each iteration. Our theoretical analysis shows that FedLaAvg achieves guaranteed convergence and a sublinear speedup with respect to the total number of clients. We implement FedLaAvg along with several baselines and evaluate them over the benchmarking MNIST and Sentiment140 data sets. The evaluation results demonstrate that FedLaAvg achieves more stable training than FedAvg in both convex and nonconvex settings and reaches a sublinear speedup. Source code and online supplement are available at the IJOC GitHub site ( http://dx.doi.org/10.1287/ijoc.2022.0057.cd , https://github.com/INFORMSJoC/2022.0057 ). History: Accepted by Ram Ramesh, Area Editor for Data Science & Machine Leaning. Funding: This work was supported by the National Key R&D Program of China [Grant 2022ZD0119100], the National Natural Science Foundation of China (NSFC) [Grants 61972252, 61972254, 62072303, 62025204, 62132018, 62202296, and 62202297], the Alibaba Innovation Research (AIR) Program, and the Tencent Rhino Bird Key Research Project. Supplemental Material: The software that supports the findings of this study is available within the paper and its Supplemental Information ( https://pubsonline.informs.org/doi/suppl/10.1287/ijoc.2022.0057 ) as well as from the IJOC GitHub software repository ( https://github.com/INFORMSJoC/2022.0057 ). The complete IJOC Software and Data Repository is available at https://informsjoc.github.io/ .
Yikai Yan, Chaoyue Niu, Zhenzhe Zheng 0001, Shaojie Tang 0001, Qinya Li, Fan Wu 0006, Chengfei Lyu, Yang-He Feng, Guihai Chen
INFORMS J. Comput.2
2024 Active Client Selection for Clustered Federated Learning
abstract
Federated learning (FL) is an emerging distributed machine learning (ML) framework that operates under privacy and communication constraints. To mitigate the data heterogeneity underlying FL, clustered FL (CFL) was proposed to learn customized models for different client groups. However, due to the lack of effective client selection strategies, the CFL process is relatively slow, and the model performance is also limited in the presence of nonindependent and identically distributed (non-IID) client data. In this work, for the first time, we propose selecting participating clients for each cluster with active learning (AL) and call our method active client selection for CFL (ACFL). More specifically, in each ACFL round, each cluster filters out a small set of clients, which are the most informative clients according to some AL metrics [e.g., uncertainty sampling, query-by-committee (QBC), loss], and aggregates only its model updates to update the cluster-specific model. We empirically evaluate our ACFL approach on the public MNIST, CIFAR-10, and LEAF synthetic datasets with class-imbalanced settings. Compared with several FL and CFL baselines, the results reveal that ACFL can dramatically speed up the learning process while requiring less client participation and significantly improving model accuracy with a relatively low communication overhead.
Honglan Huang, Yang-He Feng, Chaoyue Niu, Guangquan Cheng, Jincai Huang 0001, Zhong Liu 0002
IEEE Trans. Neural Networks Learn. Syst.4
2023 KVSAgg: Secure Aggregation of Distributed Key-Value Sets
abstract
In global data analysis, the central server needs the global statistic of the user data stored in local clients. In such cases, an Honest-but-Curious central server might put user privacy at risk in trying to collect individual statistics of each user. In response, the secure aggregation provides a solution for calculating global statistics without revealing users’ privacy data. However, existing secure aggregation protocols only focus on the data in the form of vectors or common sets, which limits their application scope. We formalize a general problem—key-value set secure aggregation—that not only includes secure vector aggregation and private set union but also supports more applications. To address the proposed problem, we devise our solution (called the KVSAgg framework) that promises satisfactory performance in security, efficiency, and accuracy. Our key technique is a homomorphic transform algorithm (called HyperIBLT) that is not only capable of bidirectionally transforming data between key-value sets and vectors, but also able to transform sum operation of sets to addition of vectors. We implement KVSAgg on both CPU and GPU platforms and perform the evaluation on three use cases including federated learning, distributed data counting, and finding global hot items. Compared with our baselines, KVSAgg simultaneously achieves the best security, efficiency higher by orders of magnitude, and zero-error in nearly all cases. All codes are open-source anonymously.
Yuhan Wu 0001, Siyuan Dong, Yikai Zhao 0001, Fangcheng Fu, Tong Yang 0003, Chaoyue Niu, Fan Wu 0006, Bin Cui 0001
ICDE7
2023 Device-Unimodal Cloud-Multimodal Collaboration for Livestreaming Content Understanding
abstract
Mobile livestreaming has revolutionized the online shopping paradigm, enabling streamers to promote products to consumers with an immersive and interactive experience. To guide consumers to the livestreams that involve their interested products, it is necessary to have a good understanding of livestreaming contents with low latency, and the key task is to accurately recognize the products being promoted by the streamers. However, the mainstream cloud-based service framework is challenged by the high concurrency of service requests, the high overhead of multimodal recognition, and the requirement of low response latency. To break the bottleneck, we propose a new device-cloud collaborative learning framework, where each streamer’s mobile device holds a unimodal recognition model that can process most of frames and also uploads the extracted unimodal features to facilitate the cloud-side multimodal recognition of the remaining few frames. In addition, the on-device unimodal model is incrementally trained over the samples constructed by leveraging the streamers’ manual labeling behaviors, thereby adapting to the heterogeneous and dynamic livestreaming contents of different streamers. Nevertheless, the device-side personalized unimodal features are misaligned in feature space and cannot be directly fused into the cloud-side multimodal model. We thus design a pluggable prompt generation module to transform the personalized unimodal features into prompt embeddings, instructing the multimodal backbone network in feature fusion. Both offline and online evaluation results reveal the effectiveness and efficiency of our design as well as its consistent advantage over existing baselines.
Chaoyue Niu, Yikai Yan, Zhijie Cao, Chengfei Lyu, Shaojie Tang 0001, Fan Wu 0006
ICDM2
2022 On-Device Learning for Model Personalization with Large-Scale Cloud-Coordinated Domain Adaption
abstract
Cloud-based learning is currently the mainstream in both academia and industry. However, the global data distribution, as a mixture of all the users' data distributions, for training a global model may deviate from each user's local distribution for inference, making the global model non-optimal for each individual user. To mitigate distribution discrepancy, on-device training over local data for model personalization is a potential solution, but suffers from serious overfitting. In this work, we propose a new device-cloud collaborative learning framework under the paradigm of domain adaption, called MPDA, to break the dilemmas of purely cloud-based learning and on-device training. From the perspective of a certain user, the general idea of MPDA is to retrieve some similar data from the cloud's global pool, which functions as large-scale source domains, to augment the user's local data as the target domain. The key principle of choosing which outside data depends on whether the model trained over these data can generalize well over the local data. We theoretically analyze that MPDA can reduce distribution discrepancy and overfitting risk. We also extensively evaluate over the public MovieLens 20M and Amazon Electronics datasets, as well as an industrial dataset collected from Mobile Taobao over a period of 30 days. We finally build a device-tunnel-cloud system pipeline, deploy MPDA in the icon area of Mobile Taobao for click-through rate prediction, and conduct online A/B testing. Both offline and online results demonstrate that MPDA outperforms the baselines of cloud-based learning and on-device training only over local data, from multiple offline and online metrics.
Yikai Yan, Chaoyue Niu, Renjie Gu, Fan Wu 0006, Shaojie Tang 0001, Lifeng Hua, Chengfei Lyu, Guihai Chen
KDD2
2022 Federated Submodel Optimization for Hot and Cold Data Features
abstract
We focus on federated learning in practical recommender systems and natural language processing scenarios. The global model for federated optimization typically contains a large and sparse embedding layer, while each client’s local data tend to interact with part of features, updating only a small submodel with the feature-related embedding vectors. We identify a new and important issue that distinct data features normally involve different numbers of clients, generating the differentiation of hot and cold features. We further reveal that the classical federated averaging algorithm (FedAvg) or its variants, which randomly selects clients to participate and uniformly averages their submodel updates, will be severely slowed down, because different parameters of the global model are optimized at different speeds. More specifically, the model parameters related to hot (resp., cold) features will be updated quickly (resp., slowly). We thus propose federated submodel averaging (FedSubAvg), which introduces the number of feature-related clients as the metric of feature heat to correct the aggregation of submodel updates. We prove that due to the dispersion of feature heat, the global objective is ill-conditioned, and FedSubAvg works as a suitable diagonal preconditioner. We also rigorously analyze FedSubAvg’s convergence rate to stationary points. We finally evaluate FedSubAvg over several public and industrial datasets. The evaluation results demonstrate that FedSubAvg significantly outperforms FedAvg and its variants.
Chaoyue Niu, Fan Wu 0006, Shaojie Tang 0001, Chengfei Lyu, Yang-He Feng, Guihai Chen
NeurIPS2
2022 Walle: An End-to-End, General-Purpose, and Large-Scale Production System for Device-Cloud Collaborative Machine Learning
Chengfei Lv, Chaoyue Niu, Renjie Gu, Xiaotang Jiang, Zhaode Wang, Ziqi Wu, Qiulin Yao, Congyu Huang, Panos Huang, Hui Shu, Jinde Song, Peng Lan, Guohuan Xu, Fei Wu 0001, Shaojie Tang 0001, Fan Wu 0006, Guihai Chen
OSDI2
2022 Pricing GAN-based data generators under Rényi differential privacy
abstract
As smart devices are becoming increasingly common in people’s daily lives, privacy and security concerns make data collection expensive and limited, which further hinder the development of data-driven tasks. This paper studies how to better conduct private data trading via a novel generator method rather than direct trading of raw data. This new method facilitates more convenient data transactions by generator, protects the privacy of data owners and is satisfactory in terms of privacy compensation and query pricing. In detail, we propose RARIEA, a market framework for tRading privAte data geneRators based on GAN under rényI diffErential privAcy, which involves data owners, a data broker, and data consumers. To start, the broker employs the GAN training generator to augment the data to relieve the data shortage, introducing noise into its training process to preserve the owners’ privacy. After that, the broker uses rényi differential privacy to quantify the privacy loss at the data item level during the GAN training process and compensates each owner according to their respective privacy policies. Finally, the data broker charges each of the data consumers for their queries, where the price is lower bounded by the total privacy compensation. We then evaluate the performance of RARIEA on classic data sets: MNIST, Fashion-MNIST, and CelebA. The analysis and simulation results reveal that the generator provided by RARIEA can not only meet the data consumers’ demand for quantity and quality but also protect the owners’ privacy. In addition, RARIEA not only allows finer control over data owner compensation, but also excels at controlling the data broker’s revenue to improve market efficiency while ensuring fairness, balance, and monotonicity of pricing.
Xikun Jiang, Chaoyue Niu, Chenhao Ying 0001, Fan Wu 0006, Yuan Luo 0003
Inf. Sci.2
2022 Toward Verifiable and Privacy Preserving Machine Learning Prediction
abstract
The ubiquitous needs for extracting insights from data are driving the emergence of service providers to offer predictions given the inputs from customers. During this process, it is important and highly nontrivial for the service providers to generate proofs of honest predictions without leaking the key parameters of their trained models. In addition, the customers are usually unwilling to reveal their sensitive inputs. In this article, we proposed MVP, which enablesMachine learning prediction in aVerifiable andPrivacy preserving fashion. MVP features the properties of polynomial decomposition and prime-order bilinear groups to simultaneously facilitate oblivious evaluation and batch outcome verification while maintaining function privacy and input privacy. We further instantiated MVP with Support Vector Machines (SVMs) and extensively evaluated its performance for the spam detection task on three practical Short Message Service (SMS) datasets. Our analysis and evaluation results reveal that MVP achieves the desired properties while incurring low computation and communication overhead.
Chaoyue Niu, Fan Wu 0006, Shaojie Tang 0001, Shuai Ma 0001, Guihai Chen
IEEE Trans. Dependable Secur. Comput.1
2022 Online Pricing With Reserve Price Constraint for Personal Data Markets
abstract
The society’s insatiable appetites for personal data are driving the emergence of data markets, allowing data consumers to launch customized queries over the datasets collected by a data broker from data owners. In this paper, we study how the data broker can maximize its cumulative revenue by posting reasonable prices for sequential queries. We thus propose a contextual dynamic pricing mechanism with the reserve price constraint, which features the properties of ellipsoid for efficient online optimization and can support linear and non-linear market value models with uncertainty. In particular, under low uncertainty, the proposed pricing mechanism attains a worst-case cumulative regret logarithmic in the number of queries. We further extend our approach to support other similar application scenarios, including hospitality service and online advertising, and extensively evaluate all three use cases over MovieLens 20M dataset, Airbnb listings in U.S. major cities, and Avazu mobile ad click dataset, respectively. The analysis and evaluation results reveal that: (1) our pricing mechanism incurs low practical regret, while the latency and memory overhead incurred is low enough for online applications; and (2) the existence of reserve price can mitigate the cold-start problem in a posted price mechanism, thereby reducing the cumulative regret.
Chaoyue Niu, Zhenzhe Zheng 0001, Fan Wu 0006, Shaojie Tang 0001, Guihai Chen
IEEE Trans. Knowl. Data Eng.1
2021 Toward Understanding the Influence of Individual Clients in Federated Learning
abstract
Federated learning allows mobile clients to jointly train a global model without sending their private data to a central server. Extensive works have studied the performance guarantee of the global model, however, it is still unclear how each individual client influences the collaborative training process. In this work, we defined a new notion, called {\em Fed-Influence}, to quantify this influence over the model parameters, and proposed an effective and efficient algorithm to estimate this metric. In particular, our design satisfies several desirable properties: (1) it requires neither retraining nor retracing, adding only linear computational overhead to clients and the server; (2) it strictly maintains the tenets of federated learning, without revealing any client's local private data; and (3) it works well on both convex and non-convex loss functions, and does not require the final model to be optimal. Empirical results on a synthetic dataset and the FEMNIST dataset demonstrate that our estimation method can approximate Fed-Influence with small bias. Further, we show an application of Fed-Influence in model debugging.
Yihao Xue, Chaoyue Niu, Zhenzhe Zheng 0001, Shaojie Tang 0001, Chengfei Lyu, Fan Wu 0006, Guihai Chen
AAAI2
2021 ERATO: Trading Noisy Aggregate Statistics over Private Correlated Data
abstract
With the commoditization of personal privacy, pricing private data has become an intriguing problem. In this paper, we study noisy aggregate statistics trading from the perspective of a data broker in data markets. We thus propose ERATO, which enables aggrEgate statistics pRicing over privATe cOrrelated data. On one hand, ERATO guarantees arbitrage freeness against cunning data consumers. On the other hand, ERATO compensates data owners for their privacy losses using both bottom-up and top-down designs. We further apply ERATO to three practical aggregate statistics, namely weighted sum, probability distribution fitting, and degree distribution, and extensively evaluate their performances on MovieLens dataset, 2009 RECS dataset, and two SNAP large social network datasets, respectively. Our analysis and evaluation results reveal that ERATO well balances utility and privacy, achieves arbitrage freeness, and compensates data owners more fairly than differential privacy based approaches.
Chaoyue Niu, Zhenzhe Zheng 0001, Fan Wu 0006, Shaojie Tang 0001, Xiaofeng Gao 0001, Guihai Chen
IEEE Trans. Knowl. Data Eng.1
2020 Online Pricing with Reserve Price Constraint for Personal Data Markets
abstract
The society's insatiable appetites for personal data are driving the emergency of data markets, allowing data consumers to launch customized queries over the datasets collected by a data broker from data owners. In this paper, we study how the data broker can maximize her cumulative revenue by posting reasonable prices for sequential queries. We thus propose a contextual dynamic pricing mechanism with the reserve price constraint, which features the properties of ellipsoid for efficient online optimization, and can support linear and non-linear market value models with uncertainty. In particular, under low uncertainty, our pricing mechanism provides a worst-case regret logarithmic in the number of queries. We further extend to other similar application scenarios, including hospitality service and online advertising, and extensively evaluate all three application instances over MovieLens 20M dataset, Airbnb listings in U.S. major cities, and Avazu mobile ad click dataset, respectively. The analysis and evaluation results reveal that our proposed pricing mechanism incurs low practical regret, online latency, and memory overhead, and also demonstrate that the existence of reserve price can mitigate the cold-start problem in a posted price mechanism, and thus can reduce the cumulative regret.
Chaoyue Niu, Zhenzhe Zheng 0001, Fan Wu 0006, Shaojie Tang 0001, Guihai Chen
ICDE1
2020 Low-viewpoint forest depth dataset for sparse rover swarms
abstract
Rapid progress in embedded computing hardware increasingly enables on-board image processing on small robots. This development opens the path to replacing costly sensors with sophisticated computer vision techniques. A case in point is the prediction of scene depth information from a monocular camera for autonomous navigation. Motivated by the aim to develop a robot swarm suitable for sensing, monitoring, and search applications in forests, we have collected a set of RGB images and corresponding depth maps. Over 100000 RGB/depth image pairs were recorded with a custom rig from the perspective of a small ground rover moving through a forest. Taken under different weather and lighting conditions, the images include scenes with grass, bushes, standing and fallen trees, tree branches, leaves, and dirt. In addition GPS, IMU, and wheel encoder data were recorded. From the calibrated, synchronized, aligned and timestamped frames about 9700 image-depth map pairs were selected for sharpness and variety. We provide this dataset to the community to fill a need identified in our own research and hope it will accelerate progress in robots navigating the challenging forest environment. This paper describes our custom hardware and methodology to collect the data, subsequent processing and quality of the data, and how to access it.
Chaoyue Niu, Danesh Tarapore, Klaus-Peter Zauner
IROS1
2020 Billion-scale federated learning on mobile clients: a submodel design with tunable privacy
abstract
Federated learning was proposed with an intriguing vision of achieving collaborative machine learning among numerous clients without uploading their private data to a cloud server. However, the conventional framework requires each client to leverage the full model for learning, which can be prohibitively inefficient for large-scale learning tasks and resource-constrained mobile devices. Thus, we proposed a submodel framework, where clients download only the needed parts of the full model, namely, submodels, and then upload the submodel updates. Nevertheless, the "position" of a client's truly required submodel corresponds to its private data, while the disclosure of the true position to the cloud server during interactions inevitably breaks the tenet of federated learning. To integrate efficiency and privacy, we designed a secure federated submodel learning scheme coupled with a private set union protocol as a cornerstone. The secure scheme features the properties of randomized response, secure aggregation, and Bloom filter, and endows each client with customized plausible deniability (in terms of local differential privacy) against the position of its desired submodel, thereby protecting private data. We further instantiated the scheme with Alibaba's e-commerce recommendation, implemented a prototype system, and extensively evaluated over 30-day Taobao user data. Empirical results demonstrate the feasibility and scalability of the proposed scheme as well as its remarkable advantages over the conventional federated learning framework, from model accuracy and convergency, practical communication, computation, and storage overhead.
Chaoyue Niu, Fan Wu 0006, Shaojie Tang 0001, Lifeng Hua, Rongfei Jia, Chengfei Lv, Guihai Chen
MobiCom1
2019 Making Big Money from Small Sensors: Trading Time-Series Data under Pufferfish Privacy
abstract
With the commoditization of personal data, pricing privacy has become an intriguing topic. In this paper, we study time-series data trading from the perspective of a data broker in data markets. We thus propose HORAE, which is a PufferfisH privacy based framewOrk for tRAding timE-series data. HORAE first employs Pufferfish privacy to quantity privacy losses under temporal correlations, and compensates data owners with distinct privacy strategies in a satisfying way. Besides, HORAE not only guarantees good profitability at the data broker, but also ensures arbitrage freeness against cunning data consumers. We further apply HORAE to physical activity monitoring, and extensively evaluate its performance on the real-world Activity Recognition with Ambient Sensing (ARAS) dataset. Our analysis and evaluation results reveal that HORAE compensates data owners in a more fine-grained manner than entry/group differential privacy based approaches, well controls the profit ratio of the data broker, and thwarts arbitrage attacks launched by data consumers.
Chaoyue Niu, Zhenzhe Zheng 0001, Shaojie Tang 0001, Xiaofeng Gao 0001, Fan Wu 0006
INFOCOM1
2019 Achieving Data Truthfulness and Privacy Preservation in Data Markets
abstract
As a significant business paradigm, many online information platforms have emerged to satisfy society's needs for person-specific data, where a service provider collects raw data from data contributors, and then offers value-added data services to data consumers. However, in the data trading layer, the data consumers face a pressing problem, i.e., how to verify whether the service provider has truthfully collected and processed data? Furthermore, the data contributors are usually unwilling to reveal their sensitive personal data and real identities to the data consumers. In this paper, we propose TPDM, which efficiently integrates Truthfulness and Privacy preservation in Data Markets. TPDM is structured internally in an Encrypt-then-Sign fashion, using partially homomorphic encryption and identity-based signature. It simultaneously facilitates batch verification, data processing, and outcome verification, while maintaining identity preservation and data confidentiality. We also instantiate TPDM with a profile matching service and a data distribution service, and extensively evaluate their performances on Yahoo! Music ratings dataset and 2009 RECS dataset, respectively. Our analysis and evaluation results reveal that TPDM achieves several desirable properties, while incurring low computation and communication overheads when supporting large-scale data markets.
Chaoyue Niu, Zhenzhe Zheng 0001, Fan Wu 0006, Xiaofeng Gao 0001, Guihai Chen
IEEE Trans. Knowl. Data Eng.1
2018 Unlocking the Value of Privacy: Trading Aggregate Statistics over Private Correlated Data
abstract
With the commoditization of personal privacy, pricing private data has become an intriguing problem. In this paper, we study noisy aggregate statistics trading from the perspective of a data broker in data markets. We thus propose ERATO, which enables aggrEgate statistics pRicing over privATe cOrrelated data. On one hand, ERATO guarantees arbitrage freeness against cunning data consumers. On the other hand, ERATO compensates data owners for their privacy losses using both bottom-up and top-down designs. We further apply ERATO to three practical aggregate statistics, namely weighted sum, probability distribution fitting, and degree distribution, and extensively evaluate their performances on MovieLens dataset, 2009 RECS dataset, and two SNAP large social network datasets, respectively. Our analysis and evaluation results reveal that ERATO well balances utility and privacy, achieves arbitrage freeness, and compensates data owners more fairly than differential privacy based approaches.
Chaoyue Niu, Zhenzhe Zheng 0001, Fan Wu 0006, Shaojie Tang 0001, Xiaofeng Gao 0001, Guihai Chen
KDD1
2017 Trading Data in Good Faith: Integrating Truthfulness and Privacy Preservation in Data Markets
abstract
As a significant business paradigm, many online information platforms have emerged to satisfy society's needs for person-specific data, where a service provider collects raw data from data contributors, and then offers value-added data services to data consumers. However, in the data trading layer, the data consumers face a pressing problem, i.e., how to verify whether the service provider has truthfully collected and processed data? Furthermore, the data contributors are usually unwilling to reveal their sensitive personal data and real identities to the data consumers. In this paper, we propose TPDM, which efficiently integrates Truthfulness and Privacy preservation in Data Markets. TPDM is structured internally in an Encrypt-then-Sign fashion, using somewhat homomorphic encryption and identitybased signature. It simultaneously facilitates batch verification, data processing, and outcome verification, while maintaining identity preservation and data confidentiality. We also instantiate TPDM with a profile-matching service, and extensively evaluate its performance on Yahoo! Music ratings dataset. Our evaluation results show that TPDM achieves several desirable properties, while incurring low computation and communication overheads when supporting a large-scale data market.
Chaoyue Niu, Zhenzhe Zheng 0001, Fan Wu 0006, Xiaofeng Gao 0001, Guihai Chen
ICDE1
2017 ERA: Towards privacy preservation and verifiability for online ad exchanges
Chaoyue Niu, Minping Zhou, Zhenzhe Zheng 0001, Fan Wu 0006, Guihai Chen
J. Netw. Comput. Appl.1
2015 An Efficient, Privacy-Preserving, and Verifiable Online Auction Mechanism for Ad Exchanges
abstract
Ad exchanges are kind of the most popular online advertising marketplaces for trading ad spaces over the Internet. Ad exchanges run auctions to sell the ad spaces on publishers' web-pages to advertisers, who want to display ads on the ad spaces. However, the parties in the auction cannot check whether the auction is carried out correctly or not. Furthermore, the advertisers are usually not willing to reveal their sensitive information when participating in the auction. In this paper, we jointly consider the auction verifiability and advertisers' privacy preservation, and propose ERA, which is an Efficient, pRivacy-preserving, and verifiAble online auction mechanism for ad exchanges. ERA exploits an Order Preserving Encryption Scheme (OPES) to guarantee privacy-preservation, and achieves verifiability by integrating a Certified Bulletin Board (CBB) and a protocol of Privacy-Preserving Integer Comparison (PPIC), which is based on the Paillier's Homomorphic Encryption Scheme (PHES). We extensively evaluate the performance of ERA, and our evaluation results show that ERA satisfies the properties of verifiability and privacy-preservation with low overhead, so ERA can be easily deployed in today's ad exchanges.
Minping Zhou, Chaoyue Niu, Zhenzhe Zheng 0001, Fan Wu 0006, Guihai Chen
GLOBECOM2