VLDB 2026 Research / reviewers in the wild / expert
Xiaopeng Ke
dblp:289/8452
· DBLP profile ↗
8ranked-venue papers
2as first author
7since 2021 · last 2025
0009-0006-0039-4013ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 2 · 1 since 2021Security and privacy · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | AQuilt: Weaving Logic and Self-Inspection into Low-Cost, High-Relevance Data Synthesis for Specialist LLMsabstractDespite the impressive performance of large language models (LLMs) in general domains, they often underperform in specialized domains.Existing approaches typically rely on data synthesis methods and yield promising results by using unlabeled data to capture domain-specific features.However, these methods either incur high computational costs or suffer from performance limitations, while also demonstrating insufficient generalization across different tasks.To address these challenges, we propose AQuilt, a framework for constructing instruction-tuning data for any specialized domains from corresponding unlabeled data, including Answer, Question, Unlabeled data, Inspection, Logic, and Task type.By incorporating logic and inspection, we encourage reasoning processes and self-inspection to enhance model performance.Moreover, customizable task instructions enable high-quality data generation for any task.As a result, we construct a dataset of 703k examples to train a powerful data synthesis model.Experiments show that AQuilt is comparable to DeepSeek-V3 while utilizing just 17% of the production cost.Further analysis demonstrates that our generated data exhibits higher relevance to downstream tasks. Xiaopeng Ke, Hexuan Deng, Xuebo Liu 0002, Jun Rao, Zhenxi Song, Jun Yu 0002, Min Zhang 0005 |
EMNLP | 1 |
| 2025 | DeepVMUnProtect: Neural Network-Based Recovery of VM-Protected Android Apps for Semantics-Aware Malware DetectionabstractThe emerging virtual machine-based Android packers render existing unpacking techniques ineffective. The state-of-the-art unpacker falls short because it relies on unreliable heuristics and manually crafted semantic models. Hence, it cannot precisely recover app semantics necessary for malware detection. In this paper, we proposeDeepVMUnProtect, a deep learning-based approach to automatically and accurately capture the semantics of VM-packed code, so as to facilitate semantic-based Android malware classification. Experiments have shown thatDeepVMUnProtectoutperforms the state-of-the-art tool on recovering opcode semantics in Qihoo(58.3%), Baidu(47.5%) and NMMP (58.8%) respectively, and can enable semantics-aware malware detection which prior work fails to do. Mu Zhang 0001, Xiaopeng Ke, Yue Duan, Sheng Zhong 0002, Fengyuan Xu |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2024 | REDLC: Learning-driven Reverse Engineering for Deep Learning CompilersabstractDeep Learning (DL) compilers such as TVM enable the efficient deployment of diverse DL models on heterogeneous and resource-constrained devices to meet the needs for low latency, privacy protection, and enhanced reliability. However, the booming of on-device DL technology will inevitably attract new types of cybercriminals and industrial spies aiming to steal commercial models. Emerging research focused on model-stealing attacks from the perspective of DL compilers mainly uses heuristic approaches, which do not work well with compiler-optimized models. This work proposes an advanced model-stealing attack pipeline that combines code representation learning and binary analysis to efficiently reverse retrainable DL framework models from TVM-compiled executables. To further improve the accuracy of reversed models, we exploit the computational relationships to correct the prediction of operators in the models using Graph Convolutional Networks. Extensive experiments demonstrate that our approach can recover 18 common DL models with different scales downloaded from Keras repositories with 99% accuracy. Yang Li 0103, Xiaopeng Ke, Fengyuan Xu, Liming Fang 0001 |
ISSRE | 4 |
| 2024 | TIM: Enabling Large-Scale White-Box Testing on In-App Deep Learning ModelsabstractIntelligent Applications (iApps), equipped with in-App deep learning (DL) models, are emerging to provide reliable DL inference services. However, in-App DL models are typically compiled into inference-only versions to enhance system performance, thereby impeding the evaluation of DL models. Specifically, the assessment of in-App models currently relies on black-box testing methods rather than direct white-box testing approaches. In this work, we propose TIM, an automated tool designed for conducting large-scale white-box testing of in-App models. Taking an iApp as input, TIM can lift the black-box (i.e., inference-only) in-App DL model into a backpropagation-enabled one and package it together, allowing comprehensive DL model testing or security issues detection. TIM proposes two reconstruction techniques to convert the inference-only model to a backpropagation-enabled version and reconstruct the DL-related IO processing code. In our experiments, we utilize TIM to extract 100 unique commercial in-App models and convert the models to white-box models, enabling backpropagation functionality. Experimental results show that TIM’s reconstruction techniques exhibit high accuracy. We open-source our prototype and part of the experimental data on the websitehttps://zenodo.org/record/7548141. Hao Wu 0067, Yuhang Gong, Xiaopeng Ke, Hanzhong Liang, Fengyuan Xu, Yunxin Liu 0001, Sheng Zhong 0002 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2023 | GAPter: Gray-Box Data Protector for Deep Learning Inference Services at User SideabstractThe widespread deployment of Deep Learning Inference Services (DLISes) has raised people’s concerns about their data privacy being breached. Although data privacy enhancement has recently attracted a lot of attention, existing solutions all require the cooperation of service providers. Users lose control of their data when making data privacy enhancement decisions. However, it is difficult to enable the user-side control of data abuse prevention because users do not have any programming skills, deep learning knowledge, or rich computing resources. In this work, we propose a fully-automatic userside data privacy enhancement solution, GAPter, for DLISes. Given such a DLIS, GAPter can adaptively fuzz the service for a suitable enhancement strategy, with no cooperation between the DLIS provider and the user. We have implemented and comprehensively evaluated GAPter. The experimental results show that GAPter can find good balance points between privacy enhancement and user data utility. Hao Wu 0067, Xiaopeng Ke, Siyi He, Fengyuan Xu, Sheng Zhong 0002 |
ICASSP | 3 |
| 2022 | Towards Practical and Efficient Long Video SummaryabstractRecently, video summarization (VS) techniques are widely used to alleviate huge processing pressure brought by numerous long videos. However, it is hard to summarize long videos efficiently since processing hundreds of frames is still time-consuming. In this paper, we find that the Kernel Temporal Segmentation (KTS) method designed for detecting the shot boundaries in SOTA VS methods is time-consuming while handling long videos. To address this issue, we propose the Distribution-based KTS (D-KTS) by fully considering the characteristic of shot length distribution. Furthermore, we propose the Hash-based Adaptive Frame Selection (HAFS) to improve the system performance by fully taking advantage of the temporal locality of long videos. Our experiments present that the proposed D-KTS is 92.70% faster and takes up 90.08% less memory than the baseline KTS method on average. Xiaopeng Ke, Boyu Chang, Hao Wu 0067, Fengyuan Xu, Sheng Zhong 0002 |
ICASSP | 1 |
| 2022 | Privacy-Preserving and Robust Federated Deep Metric LearningabstractFederated learning, in contrast to traditional learning paradigms, has demonstrated its unique advantages in providing intelligence at the edge. However, existing federated learning approaches focus on the end-to-end classification tasks requiring a simple collaboration procedure where each participant can perform its local training independently. Unfortunately, there are still many tasks relying on learning the distinguishable feature metrics with respect to all the data, which is a different collaboration procedure across training participants. For example, the model for people identification has to ensure the feature representing a person is dissimilar to those representing others. To enable such federated learning for deep metrics (a.k.a federated deep metric learning) is challenging due to the data privacy and procedure robustness issues. With the consideration of these two challenges, this work proposes a novel computing framework for federated deep metric learning. This framework leverages the system-algorithm co-design to address privacy concerns via the Trusted Execution Environment (SGX enclave) and Differential Privacy mechanism. It also introduces a large-scale federated protocol which can robustly and efficiently deal with practical factors like the network fluctuation. We implement and evaluate our computing framework with two settings. One is a real-world implementation with a large number of mobile devices, while the other one is in our controllable environment for conducting experiments in various tasks. Our evaluation results show that our computing framework is able to train federated deep metric learning models with excellent scalability, data privacy preserving, and considerable accuracy even in exception conditions. Yulong Tian, Xiaopeng Ke, Zeyi Tao, Shaohua Ding, Fengyuan Xu, Qun Li 0001, Sheng Zhong 0002 |
IWQoS | 2 |
| 2020 | Efficient Architecture Paradigm for Deep Learning Inference as a ServiceabstractDeep learning (DL) inference has been broadly used and shown excellent performance in many intelligent applications. Unfortunately, the high resource consumption and training efforts of sophisticated models present obstacles for regular users to enjoy it. Thus, Deep Learning Inference as a Service (DIaaS), offering online inference services on cloud, has earned great popularity among cloud tenants who can send their DIaaS inputs via RPCs across the internal network. However, such detached architecture paradigm is inappropriate to DIaaS because the high-dimensional inputs of DIaaS consume a lot of precious internal bandwidth and the service latency of DIaaS has to be low and stable. We therefore propose a novel architecture paradigm on cloud for DIaaS in order to address the above two problems without giving up the security and maintenance benefits. We first leverage the SGX technology, a strongly-protected user space enclave, to bring DIaaS computation to its input source as close as possible, i.e. co-locating a cloud tenant and its subscribed DIaaS in the same virtual machine. When the GPU acceleration is needed, we migrate this virtual machine to any available GPU host and transparently utilize the GPU via our backend computing stack installed on it. In this way the majority of internal bandwidth is saved compared to traditional paradigm. Furthermore, we greatly improve the efficiency of the proposed architecture paradigm, from the computation and I/O perspectives, by making the entire data flow more DL-oriented. Finally, we implement a prototype system and evaluate it in real-world scenarios. The experiments show that our locality-aware architecture achieves the average single CPU (GPU) based deep learning inference time 2.84X (4.87X) less than the traditional detached architecture on average. Xiaopeng Ke, Fengyuan Xu |
IPCCC | 2 |