VLDB 2026 Research / reviewers in the wild / expert
Dongqi Cai 0001
dblp:159/3886-1
· DBLP profile ↗
12ranked-venue papers
5as first author
12since 2021 · last 2026
0000-0003-2751-2500ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Computer networks · 5 · 3 first-author · 5 since 2021Systems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Editing as Unlearning: Are Knowledge Editing Methods Strong Baselines for Large Language Model Unlearning?abstractLarge language Model (LLM) unlearning, i.e., selectively removing information from LLMs, is vital for responsible model deployment. Differently, LLM knowledge editing aims to modify LLM knowledge instead of removing it. Though editing and unlearning seem to be two distinct tasks, we find there is a tight connection between them. In this paper, we conceptualize unlearning as a special case of editing where information is modified to a refusal or "empty set" response, signifying its removal. This paper thus investigates if knowledge editing techniques are strong baselines for LLM unlearning. We evaluate state-of-the-art (SOTA) editing methods (e.g., ROME, MEMIT, GRACE, WISE, and AlphaEdit) against existing unlearning approaches on pretrained and finetuned knowledge. Results show certain editing methods, notably WISE and AlphaEdit, are effective unlearning baselines, especially for pretrained knowledge, and excel in generating human-aligned refusal answers. To better adapt editing methods for unlearning applications, we propose practical recipes including self-improvement and query merging. The former leverages the LLM's own in-context learning ability to craft a more human-aligned unlearning target, and the latter enables ROME and MEMIT to perform well in unlearning longer sample sequences. We advocate for the unlearning community to adopt SOTA editing methods as baselines and explore unlearning from an editing perspective for more holistic LLM memory control. Zexi Li 0001, Xiangzhu Wang, William F. Shen, Meghdad Kurmanji, Xinchi Qiu, Dongqi Cai 0001, Chao Wu 0001, Nicholas D. Lane |
AAAI | 6 |
| 2026 | FwdLLM+: Accelerating Forward-Only FedLLM With Low-Rank PerturbationsabstractFederated Learning (FL) facilitates privacy-preserving fine-tuning of Large Language Models (LLMs) for mobile applications, termed FedLLM. A vital challenge of FedLLM is the tension between LLM complexity and resource constraint of mobile devices. In response to this challenge, we first introduceFwdLLM(our conference version), an innovative FL framework designed to enhance the FedLLM efficiency. The key idea ofFwdLLMis to employ backpropagation (BP)-free training methods, requiring devices only to execute memory-efficient “perturbed inference”. Enabled by mobile NPU acceleration and an expanded array of participating devices,FwdLLMdelivers substantially better wall-clock efficiency than BP-based FedLLM. However,FwdLLMbased on vanilla BP-free optimization theoretically requires more optimization steps to converge. In this work, we further enhanceFwdLLMtoFwdLLM+, which incorporates advanced zeroth-order optimization techniques and low-rank perturbation decomposition to reduce convergence steps. Finally, we conduct extensive experiments on 4 models (ranging from 110 M to 7B) and 8 more datasets, demonstrating thatFwdLLM+achieves up to 151× faster training, a 93$\%$memory reduction compared to vanilla BP-based FedLLM, and superior performance compared toFwdLLM, enabling efficient federated fine-tuning of billion-parameter LLMs on commodity mobile devices. Mengwei Xu 0001, Zhenyan Lu, Wei Liu 0302, Shangguang Wang, Nicholas D. Lane, Qibo Sun, Dongqi Cai 0001 |
IEEE Trans. Mob. Comput. | 8 |
| 2025 | Demystifying Small Language Models for Edge DeploymentabstractSmall language models (SLMs) have emerged as a promising solution for deploying resource-constrained devices, such as smartphones and Web of Things. This work presents the first comprehensive study of over 60 SLMs such as Microsoft Phi and Google Gemma that are publicly accessible. Our findings show that state-of-the-art SLMs outperform 7B models in general tasks, proving their practical viability. However, SLMs’ in-context learning capabilities remain limited, and their efficiency has significant optimization potential. We identify key SLM optimization opportunities, including dynamic task-specific routing, model-hardware co-design, and vocabulary/KV cache compression. Overall, we expect the work to reveal an all-sided landscape of SLMs, benefiting the research community across algorithm, model, system, and hardware levels. Zhenyan Lu, Xiang Li 0067, Dongqi Cai 0001, Rongjie Yi, Fangming Liu, Wei Liu 0302, Jian Luan 0001, Nicholas D. Lane, Mengwei Xu 0001 |
ACL (1) | 3 |
| 2025 | DEPT: Decoupled Embeddings for Pre-training Language ModelsabstractLanguage Model pre-training uses broad data mixtures to enhance performance across domains and languages. However, training on such heterogeneous text corpora requires extensive and expensive efforts. Since these data sources vary significantly in lexical, syntactic, and semantic aspects, they cause negative interference or the ``curse of multilinguality''. To address these challenges we propose a communication-efficient pre-training framework, DEPT. Our method decouples embeddings from the transformer body while simultaneously training the latter on multiple data sources without requiring a shared vocabulary. DEPT can: (1) train robustly and effectively under significant data heterogeneity, (2) minimize token embedding parameters to only what the data source vocabulary requires, while cutting communication costs in direct proportion to both the communication frequency and the reduction in parameters, (3) enhance transformer body plasticity and generalization, improving both average perplexity (up to 20%) and downstream task performance, and (4) enable training with custom optimized vocabularies per data source. We demonstrate DEPT's potential via the first vocabulary-agnostic federated pre-training of billion-scale models, reducing communication costs by orders of magnitude and embedding memory by 4-5x. Alex Iacob, Lorenzo Sani, Meghdad Kurmanji, William F. Shen, Xinchi Qiu, Dongqi Cai 0001, Yan Gao 0016, Nicholas D. Lane |
ICLR | 6 |
| 2025 | ShortcutsBench: A Large-Scale Real-world Benchmark for API-based AgentsabstractRecent advancements in integrating large language models (LLMs) with application programming interfaces (APIs) have gained significant interest in both academia and industry. Recent work demonstrates that these API-based agents exhibit relatively strong autonomy and planning capabilities. However, their ability to handle multi-dimensional difficulty levels, diverse task types, and real-world demands remains unknown.
In this paper, we introduce \textsc{ShortcutsBench}, a large-scale benchmark for the comprehensive evaluation of API-based agents in solving real-world complex tasks. \textsc{ShortcutsBench} includes a wealth of real APIs from Apple Inc., refined user queries, human-annotated
high-quality action sequences, detailed parameter filling values, and parameters requesting necessary input from the system or user. We revealed how existing benchmarks~/~datasets struggle to accommodate the advanced reasoning capabilities of existing more intelligent LLMs. Moreover, our extensive evaluation of agents built with $5$ leading open-source (size $\geq$ 57B) and $5$ closed-source LLMs (e.g. Gemini-1.5-Pro and GPT-4o-mini) with varying intelligence level reveals significant limitations of existing API-based agents in the whole process of handling complex queries related to API selection, parameter filling, and requesting necessary input from the system and the user. These findings highlight the great challenges that API-based agents face in effectively fulfilling real and complex user queries. All datasets, code, experimental logs, and results are available at https://github.com/EachSheep/ShortcutsBench Haiyang Shen, Desong Meng, Dongqi Cai 0001, Li Zhang 0133, Mengwei Xu 0001, Yun Ma 0002 |
ICLR | 4 |
| 2025 | Federated LLM Pre-Training on Mobile PhonesabstractOver the past decades, billions of mobile phones have become the primary interface to the Internet, accessing vast amounts of private user data. These devices are idle most of the time or become obsolete after a few years. Pre-training Large Language Models (LLMs) on ubiquitous mobile phones offers a promising way to utilize both private data and idle computing power. In this work, we propose the first federated LLM pre-training framework for mobile devices and demonstrate that it can potentially achieve wall-clock training time comparable to centralized pre-training. Dongqi Cai 0001 |
MobiSys | 1 |
| 2024 | Mobile Foundation Model as FirmwareabstractIn the current AI era, mobile devices such as smartphones are tasked with executing a myriad of deep neural networks (DNNs) locally. It presents a complex landscape, as these models are highly fragmented in terms of architecture, operators, and implementations. Such fragmentation poses significant challenges to the co-optimization of hardware, systems, and algorithms for efficient and scalable mobile AI. Jinliang Yuan, Chen Yang 0043, Dongqi Cai 0001, Shihe Wang, Zeling Zhang, Xiang Li 0067, Dingge Zhang, Hanzi Mei, Xianqing Jia, Shangguang Wang, Mengwei Xu 0001 |
MobiCom | 3 |
| 2024 | SILENCE: Protecting privacy in offloaded speech understanding on resource-constrained devicesabstractSpeech serves as a ubiquitous input interface for embedded mobile devices.
Cloud-based solutions, while offering powerful speech understanding services, raise significant concerns regarding user privacy.
To address this, disentanglement-based encoders have been proposed to remove sensitive information from speech signals without compromising the speech understanding functionality.
However, these encoders demand high memory usage and computation complexity, making them impractical for resource-constrained wimpy devices.
Our solution is based on a key observation that speech understanding hinges on long-term dependency knowledge of the entire utterance, in contrast to privacy-sensitive elements that are short-term dependent.
Exploiting this observation, we propose SILENCE, a lightweight system that selectively obscuring short-term details, without damaging the long-term dependent speech understanding performance.
The crucial part of SILENCE is a differential mask generator derived from interpretable learning to
automatically configure the masking process.
We have implemented SILENCE on the STM32H7 microcontroller and evaluate its efficacy under different attacking scenarios.
Our results demonstrate that SILENCE offers speech understanding performance and privacy protection capacity comparable to existing encoders, while achieving up to 53.3$\times$ speedup and 134.1$\times$ reduction in memory footprint. Dongqi Cai 0001, Shangguang Wang, Zeling Zhang, Felix Xiaozhu Lin, Mengwei Xu 0001 |
NeurIPS | 1 |
| 2024 | FwdLLM: Efficient Federated Finetuning of Large Language Models with Perturbed Inferences
Mengwei Xu 0001, Dongqi Cai 0001, Yaozong Wu, Xiang Li 0067, Shangguang Wang |
USENIX ATC | 2 |
| 2024 | Accelerating Vertical Federated LearningabstractPrivacy, security and data governance constraints rule out a brute force process in the integration of cross-silo data, which inherits the development of the Internet of Things. Federated learning is proposed to ensure that all parties can collaboratively complete the training task while the data is not out of the local. Vertical federated learning is a specialization of federated learning for distributed features. To preserve privacy, homomorphic encryption is applied to enable encrypted operations without decryption. Nevertheless, together with a robust security guarantee, homomorphic encryption brings extra communication and computation overhead. In this paper, we analyze the current bottlenecks of vertical federated learning under homomorphic encryption comprehensively and numerically. We propose a straggler-resilient and computation-efficient accelerating system that reduces the communication overhead in heterogeneous scenarios by 65.26% at most and reduces the computation overhead caused by homomorphic encryption by 40.66% at most. Our system can improve the robustness and efficiency of the current vertical federated learning framework without loss of security. Dongqi Cai 0001, Tao Fan 0002, Yan Kang 0001, Lixin Fan, Mengwei Xu 0001, Shangguang Wang, Qiang Yang 0001 |
IEEE Trans. Big Data | 1 |
| 2023 | Efficient Federated Learning for Modern NLPabstractTransformer-based pre-trained models have revolutionized NLP for superior performance and generality. Fine-tuning pre-trained models for downstream tasks often requires private data, for which federated learning is the de-facto approach (i.e., FedNLP). However, our measurements show that FedNLP is prohibitively slow due to the large model sizes and the resultant high network/computation cost. Towards practical FedNLP, we identify as the key building blocks adapters, small bottleneck modules inserted at a variety of model layers. A key challenge is to properly configure the depth and width of adapters, to which the training speed and efficiency is highly sensitive. No silver-bullet configuration exists: the optimal choice varies across downstream NLP tasks, desired model accuracy, and mobile resources. To automate adapter configuration, we propose AdaFL1, a framework that enhances the existing FedNLP with two novel designs. First, AdaFL progressively upgrades the adapter configuration throughout a training session; the principle is to quickly learn shallow knowledge by only training fewer and smaller adapters at the model's top layers, and incrementally learn deep knowledge by incorporating deeper and larger adapters. Second, AdaFL continuously profiles future adapter configurations by allocating participant devices to trial groups. Extensive experiments show that AdaFL can reduce FedNLP's model convergence delay to no more than several hours, which is up to 155.5× faster compared to vanilla FedNLP and 48× faster compared to strong baselines. Dongqi Cai 0001, Yaozong Wu, Shangguang Wang, Felix Xiaozhu Lin, Mengwei Xu 0001 |
MobiCom | 1 |
| 2023 | Federated Few-Shot Learning for Mobile NLPabstractNatural language processing (NLP) sees rich mobile applications. To support various language understanding tasks, a foundation NLP model is often fine-tuned in a federated, privacy-preserving setting (FL). This process currently relies on at least hundreds of thousands of labeled training samples from mobile clients; yet mobile users often lack willingness or knowledge to label their data. Such an inadequacy of data labels is known as a few-shot scenario; it becomes the key blocker for mobile NLP applications. Dongqi Cai 0001, Shangguang Wang, Yaozong Wu, Felix Xiaozhu Lin, Mengwei Xu 0001 |
MobiCom | 1 |