VLDB 2026 Research / reviewers in the wild / expert
Fengyuan Xu
dblp:54/6929
· DBLP profile ↗
77ranked-venue papers
6as first author
47since 2021 · last 2026
0000-0003-3388-7544ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 30 · 4 first-author · 14 since 2021Security and privacy · 18 · 10 since 2021Databases, data management, data science and information retrieval · 9 · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 9 since 2021Systems, architecture and hardware · 7 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Diverse Human Driving Vehicle Simulation in Background Traffic for Autonomous Driving TestsabstractRealistic background traffic is critical to the simulation platforms for autonomous driving (AD) testing. Given that most vehicles in reality are driven by human beings, introducing human driving (HD) vehicles to the background traffic is necessary to be able to discover more problems of the tested AD vehicle in the simulation stage. However, existing methods rely on ad-hoc rules or data-driven training to mimic partial human driver behaviors, which are not comprehensive and lack transparency. In this work, we design a smart human driving vehicle simulator HDSim which is empowered by cognitively inspired modeling and AI models. HDSim enables diverse, realistic, and scalable HD traffic simulation on AD testing platforms like CARLA in a non-intrusive manner. There are two novel components in HDSim. First, we introduce a driver model to guide the generation of diverse human driving styles by using different combinations of latent cognitive factors in a hierarchy. Second, we design a Perception-Mediated Behavior Influence (PMBI) mechanism to use LLM-assisted perceptual transformations to indirectly fuse driving actions with driving styles. Experiments show that HDSim traffic can help simulation platforms like CARLA to reveal 68% more failures of tested AD vehicles, and the explainability of reported accidents is also improved. Wendi Li, Hao Wu 0067, Bing Mao 0001, Fengyuan Xu, Sheng Zhong 0002 |
AAAI | 5 |
| 2026 | Enhancing All-to-X Backdoor Attacks with Optimized Target Class MappingabstractBackdoor attacks pose severe threats to machine learning systems, prompting extensive research in this area. However, most existing work focuses on single-target All-to-One (A2O) attacks, overlooking the more complex All-to-X (A2X) attacks with multiple target classes, which are often assumed to have low attack success rates. In this paper, we first demonstrate that A2X attacks are robust against state-of-the-art defenses. We then propose a novel attack strategy that enhances the success rate of A2X attacks while maintaining robustness by optimizing grouping and target class assignment mechanisms. Our method improves the attack success rate by up to 28%, with average improvements of 6.7%, 16.4%, 14.1% on CIFAR10, CIFAR100, and Tiny-ImageNet, respectively. We anticipate that this study will raise awareness of A2X attacks and stimulate further research in this underexplored area. Yulong Tian, Fengyuan Xu |
AAAI | 4 |
| 2026 | KAT: Knowledge-Context Augmentation for Evolving LLM-Based Telecom Troubleshooting
Feng Lyu 0001, Hao Wu 0067, Shucheng Li, Fan Wu 0014, Fengyuan Xu |
INFOCOM | 9 |
| 2026 | Seeing the Whole Through the Parts: Discovering Objects through Semantic Part Mining in Weak Supervision
Shucheng Li, Weixuan Xu, Hao Wu 0067, Fengyuan Xu, Fan Wu 0014, Feng Lyu 0001 |
SIGIR | 5 |
| 2026 | WAMO: Toward Secure Browser Inference via Web Model Obfuscation in WebAssemblyabstractArtificial intelligence (AI) models are increasingly deployed directly in web browsers to enable low-latency, privacy-preserving inference. While this shift offers significant usability and scalability benefits, it also exposes model code and parameters to untrusted environments, leaving them vulnerable to theft, reverse engineering, and tampering. Our analysis demonstrates that existing JavaScript-based inference frameworks are highly susceptible to model extraction, posing serious security and intellectual property risks. To address this gap, we present WAMO, a WebAssembly-based obfuscation framework that secures browser-side AI models. WAMO introduces a comprehensive conversion pipeline that translates mainstream model formats into Wasm-native modules, applying model-specific obfuscation at the Wasm layer to target weights, operators, and computation graphs. This design shifts model execution from easily inspected JavaScript assets to hardened Wasm binaries, significantly raising the difficulty of static and dynamic analysis. Evaluation shows that WAMO increases cyclomatic complexity by 71.0% and Halstead effort by 455.57%, while incurring < 1% accuracy loss and no inference slowdown. Pengfei Yu 0002, Jingjing Gu, Fengyuan Xu, Xinyi Huang 0001 |
WWW | 5 |
| 2026 | H2O: Heterogeneity-Aware Hierarchical Orchestration for Memory-Efficient On-Device LLM InferenceabstractOn-device Large Language Model (LLM) inference enables private, personalized AI but faces memory constraints. Despite memory optimization efforts, scaling laws continue to increase model sizes and memory pressure. In this paper, we revisit the core memory bottlenecks in on-device LLM inference and conduct a comprehensive analysis of mainstream optimization techniques. We uncover several overlooked inefficiencies: (1) model weights, not KV caches, dominate memory usage; (2) weight sparsity remains underutilized; (3) OS-level memory behaviors cause redundancy; and (4) naive weight loading leads to excessive memory residency. To address these challenges, we propose H2O, a heterogeneity-aware hierarchical orchestration framework for memory-efficient on-device LLM inference. H2O introduces three key techniques, including hierarchical weight orchestration to reduce redundant memory retention, zero copy I/O–compute parallelism for safe and efficient memory reuse, and heterogeneity-aware inference planning to adapt to diverse mobile hardware constraints. Extensive experimental results show that H2O reduces peak memory usage by up to 60%, eliminates out-of-memory (OOM) failures for 7B–13B models, and improves inference latency by 34%–94% under tight memory budgets. We open-source our implementation at: https://github.com/ccfeiker/H2O. Feng Lyu 0001, Hao Wu 0067, Zhanxi Li, Shucheng Li, Fengyuan Xu |
IEEE Trans. Mob. Comput. | 6 |
| 2026 | SlimFit-Gens: Toward Low Bandwidth One-on-One Video Calls on COTS SmartphonesabstractMobile video calls play an essential role in our daily lives. However, in bandwidth-limited scenarios (e.g., inadequate cellular coverage, congested satellite links, and metered connections), users often experience poor quality of experience (QoE) during video calls. While recent advances in deep learning have demonstrated significant improvements in video compression over traditional methods, existing approaches are ill-suited for bidirectional video streaming on smartphones. The primary challenge lies in simultaneously achieving high video quality, computational and bandwidth efficiency, and practical usability on constrained mobile devices. In this work, we present SlimFit-Gens, the first practical video calling system for smartphones capable of delivering real-time 480p video at as low as 30 kbps. SlimFit-Gens addresses the challenge with joint algorithm and system-level optimizations. The core technique is a fine-grained model personalization design tailored for mobile video calling, enabling high-fidelity video generation at low model complexity. SlimFit-Gens achieves effective personalized adaptation through a novel two-stage personalization mechanism working upon an optimized model architecture. It also incorporates a privacy preserving, resource-efficient system design, featuring TEE-based (e.g., Confidential VM/NVIDIA Confidential Computing) fine-tuning on the server side and heterogeneity-aware inference on the device side. We implement SlimFit-Gens on four commercial off-the-shelf (COTS) smartphones with different system-on-chip (SoC) configurations and conduct extensive evaluations. Compared to prior work, SlimFit-Gens simultaneously improves generation quality with a 0.09-0.12 reduction in LPIPS and system efficiency through a 1.6-1.8× increase in video frame rate. Jingzhou Zhu, Lizhi Sun, Peiwen Dong, Wendi Li, Yixin Xu 0003, Hao Wu 0067, Fengyuan Xu, Sheng Zhong 0002 |
IEEE Trans. Mob. Comput. | 10 |
| 2025 | GET-AID: Graph-Enhanced Transformer for Provenance-Based Advanced Persistent Threats Investigation and Detection
Fengyuan Xu, Jiahong Yang 0003, Wenting Li 0002, Zonghua Zhang, Chenbin Zhang, Meng Ma 0001, Ping Wang 0003 |
ESORICS (4) | 2 |
| 2025 | Multi-level Feature Interaction and Local-global Dynamic Fusion Network for Medical Image SegmentationabstractIn recent years, U-shaped networks have been extensively explored in medical image segmentation and achieved remarkable performance. However, existing U-shaped segmentation methods fail to select the most representative information from different encoding layers for fusion with the decoder, resulting in blurred edge details. Furthermore, most segmentation methods struggle to effectively fuse the features extracted from Convolutional Neural Networks (CNNs) and Transformer networks, leading to unsatisfactory segmentation results. To address these issues, in this article, we propose a Multi-Level Feature Interaction Dynamic Fusion Network (MFIDF-Net). Specifically, a Multi-Level Feature Interaction (MLFI) method is proposed to fuse different encoders and select the most representative features from the channel, spatial and multi-scale dimensions to be passed to the decoder, thus enhancing edge detail segmentation and improving overall scene recognition performance. Furthermore, We propose a Local-Global Dynamic Fusion (LGDF) module, which adaptively generates two dynamic weights based on the contribution of CNN and Transformer features at each layer to weight and fuse their features, thereby boosting the representational ability of feature fusion and strengthening segmentation performance. The experimental results validate the superiority of MFIDF-Net in medical image segmentation on challenging benchmark datasets. Fengyuan Xu, Junqing Liu |
IJCNN | 1 |
| 2025 | Unleashing the Power of LLM to Infer State Machine From the Protocol ImplementationabstractState machines are essential for enhancing protocol analysis to identify vulnerabilities. However, inferring state machines from network protocol implementations is challenging due to complex code syntax and semantics. Traditional dynamic analysis methods often miss critical state transitions due to limited coverage, while static analysis faces path explosion issues. To overcome these challenges, we introduce a novel state machine inference approach utilizing Large Language Models (LLMs), named ProtocolGPT. This method employs retrieval augmented generation technology to enhance a pre-trained model with specific knowledge from protocol implementations. Through effective prompt engineering, we accurately identify and infer state machines. To the best of our knowledge, our approach represents the first state machine inference that leverages the source code of protocol implementations. Our evaluation of six protocol implementations shows that our method achieves a precision of over 90 %, outperforming the baselines by more than 30 %. Furthermore, integrating our approach with protocol fuzzing improves coverage by more than 20 % and uncovers two 0-day vulnerabilities compared to baseline methods. Haiyang Wei, Ligeng Chen, Zhengjie Du, Haohui Huang, Guang Cheng 0001, Fengyuan Xu, Linzhang Wang, Bing Mao 0001 |
IWQoS | 8 |
| 2025 | Training Data Attribution: Was Your Model Secretly Trained On Data Created By Mine?abstractThe emergence of text-to-image models has recently sparked significant interest, but the attendant is a looming shadow of potential infringement by violating user terms. Specifically, an adversary may exploit data created by a commercial model to train their own without proper authorization. To address such risk, it is crucial to investigate the attribution of a suspicious model's training data by determining whether its training data originates, wholly or partially, from a specific source model. To trace the generated data, existing methods need to apply additional watermarks during either the training or inference phases of the source model. However, these methods are impractical for pre-trained models that have been released, especially when model owners lack security expertise. To tackle this challenge, we propose an injection-free training data attribution method for text-to-image models. It can identify whether a model's training data stems from a certain source model without adding additional watermarks on the source model. The rationale of our method lies in the inherent memorization characteristic of text-to-image models. The memorization of training data is inherited through the data generated by the source model to the model trained on that data, making the source model and the infringing model exhibit consistent behaviors on specific samples. Therefore, from instance-level, we develop detection-based and generation-based strategies to uncover these distinct samples and using them as inherent watermarks to verify if a suspicious model originates from the source model. Besides, we also propose a statistical-level attribution method, utilizing the shadow model technique to train an attribution discriminator. Experiments demonstrate that the attribution accuracy and AUC scores of our methods are over 80% even when the infringing model only uses a small proportion of generated data. Hao Wu 0067, Lingcui Zhang, Fengyuan Xu, Jin Cao 0001, Fenghua Li 0001, Ben Niu 0001 |
KDD (2) | 4 |
| 2025 | When LLMs Copy to Think: Uncovering Copy-Guided Attacks in Reasoning LLMsabstractLarge Language Models (LLMs) have become integral to automated code analysis, enabling tasks such as vulnerability detection and code comprehension. However, their integration introduces novel attack surfaces. In this paper, we identify and investigate a new class of prompt-based attacks, termed Copy-Guided Attacks (CGA), which exploit the inherent copying tendencies of reasoning-capable LLMs. By injecting carefully crafted triggers into external code snippets, adversaries can induce the model to replicate malicious content during inference. This behavior enables two classes of vulnerabilities: inference length manipulation, where the model generates abnormally short or excessively long reasoning traces; and inference result manipulation, where the model produces misleading or incorrect conclusions. We formalize CGA as an optimization problem and propose a gradient-based approach to synthesize effective triggers. Empirical evaluation on state-of-the-art reasoning LLMs shows that CGA reliably induces infinite loops, premature termination, false refusals, and semantic distortions in code analysis tasks. While highly effective in targeted settings, we observe challenges in generalizing CGA across diverse prompts due to computational constraints, posing an open question for future research. Our findings expose a critical yet underexplored vulnerability in LLM-powered development pipelines and call for urgent advances in prompt-level defense mechanisms. Yue Li 0002, Xiao Li 0082, Hao Wu 0067, Yue Zhang 0025, Fengyuan Xu, Xiuzhen Cheng, Sheng Zhong 0002 |
MASS | 5 |
| 2025 | PriCAF: Privacy-Preserving Contribution Assessment in Federated Learning Before Model Training
Yixin Xu 0003, Hao Wu 0067, Jingzhou Zhu, Fengyuan Xu, Sheng Zhong 0002 |
ACM Multimedia | 4 |
| 2025 | Make a Feint to the East While Attacking in the West: Blinding LLM-Based Code Auditors with Flashboom AttacksabstractLLM-based vulnerability auditors (e.g., GitHub Copilot) represent a significant advancement in automated code analysis, offering precise detection of security vulnerabilities. This paper explores the potential to circumvent LLM-based vulnerability auditors by diverting their focus, decided by the LLM attention mechanism, away from real vulnerable code segments. In these LLM-based vulnerability auditors, the attention mechanism is supposed to focus on potentially vulnerable code sections to identify security issues. Our approach introduces high-attention code snippets (code fragments designed to draw focus) into the codebase under review. By strategically diverting the model's focus away from actual vulnerabilities, this technique effectively “blinds” the LLM, resulting in missed detections. To scale this approach, we present Crazy-Ivan11Source code, dataset and attack results are available at https://github.com/oxygen-hunter/Flashboom., an automated system that identifies and seamlessly integrates high-attention code snippets, shifting focus away from genuine vulnerabilities to decoy functions. Through systematic function-level prioritization and refinement, Crazy-Ivan optimizes the blinding effect, producing the Flashboom that can reduce the model's capacity to detect true security risks. Our evaluation underscores the effectiveness of Flashboom, achieving blinding success rates of up to 96.3% on CodeLlama and 83.05% on Gemma, with notable cross-model transferability and applicability across multiple programming languages. In a case study with GitHub Copilot, Flashboom led the tool to overlook a critical blockchain vulnerability, underscoring the security implications of such attention-diverting attacks and the risks inherent in relying solely on LLM-based automated auditing systems. We have reported our findings to the respective LLM-based code auditor vendors, who have acknowledged the issues and are currently working on fixes. Xiao Li 0082, Yue Li 0002, Hao Wu 0067, Yue Zhang 0025, Kaidi Xu, Xiuzhen Cheng, Sheng Zhong 0002, Fengyuan Xu |
SP | 8 |
| 2025 | OPRE: Towards Better Availability of PCNs Through RecoveringabstractThe Payment Channel Network (PCN) stands out as one of the most promising technologies for scaling blockchain-based cryptocurrencies. However, a noteworthy challenge arises during the utilization of PCNs, where a substantial portion of payment channels gradually becomes exhausted, leading to a reduction in the overall availability of PCNs. This issue is crucial in the context of blockchain off-chain PCNs and warrants a comprehensive investigation. In this paper, we introduce the problem of optimal recover and propose OPtimal REcovering protocols, denoted asOPREandOPRE+, to address this challenge. The protocols target at recovering the optimal number of nearly exhausted channels in the PCN. OPRE provides a basic solution, and OPRE+ is an augmentation which provides a more efficient and effective solution. Furthermore, to address users’ privacy concerns, we propose privacy-preserving versions of the protocols, ensuring that users’ balance on payment channels remains undisclosed during the execution of the protocols. Beyond the theoretical design and analysis, we implement these protocols and conduct experimental evaluations to assess their performance. The results affirm that our protocols exhibit efficiency and effectiveness in significantly improving the availability of PCNs. Minze Xu, Yue Li 0002, Chenglu Shi, Yuan Zhang 0004, Yongchuan Niu, Fengyuan Xu, Sheng Zhong 0002 |
IEEE Trans. Dependable Secur. Comput. | 6 |
| 2025 | UTRDCL: Stealthy DCL-Based Obfuscation and Its Attacks and Defenses in AndroidabstractDynamic Class Loading (DCL) is a legitimate technique extensively used by Android developers to incorporate additional functionalities into applications at runtime. However, adversaries can exploit DCL as a stealthy obfuscation technique to dynamically load malicious code and evade detection. While prior studies have analyzed typical DCL-based obfuscation and the attacks it enables—such as identifying payloads on storage, inspecting DCL-related APIs, or profiling dynamic behaviors—existing solutions remain insufficient against increasingly evasive DCL threats. In this paper, we propose UTRDCL, a novel stealthy obfuscation technique that leverages system APIs instead of conventional DCL-related APIs, and employs an automated footprint cleanup strategy to minimize runtime traces. Based on UTRDCL, we construct three real-world attack instances by embedding it into existing malware and benign applications, demonstrating how it can be used to evade detection. To counter such threats, we design and implement a lightweight defense mechanism by patching a previously overlooked vulnerability in the Android system that UTRDCL exploits in this specific context. This system-level mitigation closes the attack surface leveraged by UTRDCL, offering a more fundamental defense than behavioral detection. Extensive experiments show that attacks leveraging UTRDCL can evade 11 state-of-the-art malware detectors from open-source, academic, and commercial sources. We further validate our defense mechanism on real devices, demonstrating its effectiveness in preventing UTRDCL-based attacks without introducing noticeable overhead. Our proof-of-concept of UTRDCL and its defense is publicly available1. Hao Wu 0067, Sheng Zhong 0002, Fengyuan Xu |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2025 | DeepVMUnProtect: Neural Network-Based Recovery of VM-Protected Android Apps for Semantics-Aware Malware DetectionabstractThe emerging virtual machine-based Android packers render existing unpacking techniques ineffective. The state-of-the-art unpacker falls short because it relies on unreliable heuristics and manually crafted semantic models. Hence, it cannot precisely recover app semantics necessary for malware detection. In this paper, we proposeDeepVMUnProtect, a deep learning-based approach to automatically and accurately capture the semantics of VM-packed code, so as to facilitate semantic-based Android malware classification. Experiments have shown thatDeepVMUnProtectoutperforms the state-of-the-art tool on recovering opcode semantics in Qihoo(58.3%), Baidu(47.5%) and NMMP (58.8%) respectively, and can enable semantics-aware malware detection which prior work fails to do. Mu Zhang 0001, Xiaopeng Ke, Yue Duan, Sheng Zhong 0002, Fengyuan Xu |
IEEE Trans. Inf. Forensics Secur. | 7 |
| 2024 | MuSR: Multi-Scale 3D Scenes Reconstruction based on Monocular VideoabstractThree-dimensional (3D) scene reconstruction, particularly from monocular videos, is a significant challenge in large-scale scenarios due to difficulty handling varying object sizes and high computational resource needs. This paper introduces MuSR, a novel multi-scale reconstruction method addressing these issues. MuSR features a dynamic multi-resolution spatial structure that adaptively adjusts voxel resolution for objects of different sizes to improve reconstruction quality. MuSR also employs a block-based sparse 3D data structure and hardware resource management strategy to reduce GPU memory usage while maintaining efficient reconstruction. Evaluated on ScanNet and 7-Scenes datasets, as well as real-world scenes, MuSR outperforms state-of-the-art methods in terms of efficiency, completeness, and geometric shape reconstruction, proving its applicability in practical multi-scale 3D reconstructions. Hao Wu 0067, Peiwen Dong, Yixin Xu 0003, Fengyuan Xu, Sheng Zhong 0002 |
ICASSP | 5 |
| 2024 | VIDAR: Data Quality Improvement for Monocular 3D Reconstruction through In-situ Visual Interactionabstract3D reconstruction based on monocular videos has attracted wide attention, and existing reconstruction methods usually work in a reconstruction-after-scanning manner. However, these methods suffer from insufficient data collection problems due to the lack of effective guidance for users during the scanning process, which affects reconstruction quality. We propose VIDAR, which visually guides users with the streaming incremental reconstructed mesh in data collection for monocular 3D reconstruction. We propose an incremental mesh extraction algorithm to achieve lossless fusion of streaming incremental mesh data via slice-style management for guidance quality. We also design an incremental mesh rendering algorithm to achieve precise memory reallocation by updating the buffer in a fill-in-the-blank pattern for guidance efficiency. Besides, we introduce several optimizations on data transmission and human-computer interaction to improve the overall system performance. The experiment results on real-world scenes show that VIDAR efficiently delivers high-quality visual guidance and outperforms the non-interactive data collection methods for scene reconstruction. Hao Wu 0067, Fengyuan Xu, Sheng Zhong 0002 |
ICRA | 5 |
| 2024 | REDLC: Learning-driven Reverse Engineering for Deep Learning CompilersabstractDeep Learning (DL) compilers such as TVM enable the efficient deployment of diverse DL models on heterogeneous and resource-constrained devices to meet the needs for low latency, privacy protection, and enhanced reliability. However, the booming of on-device DL technology will inevitably attract new types of cybercriminals and industrial spies aiming to steal commercial models. Emerging research focused on model-stealing attacks from the perspective of DL compilers mainly uses heuristic approaches, which do not work well with compiler-optimized models. This work proposes an advanced model-stealing attack pipeline that combines code representation learning and binary analysis to efficiently reverse retrainable DL framework models from TVM-compiled executables. To further improve the accuracy of reversed models, we exploit the computational relationships to correct the prediction of operators in the models using Graph Convolutional Networks. Extensive experiments demonstrate that our approach can recover 18 common DL models with different scales downloaded from Keras repositories with 99% accuracy. Yang Li 0103, Xiaopeng Ke, Fengyuan Xu, Liming Fang 0001 |
ISSRE | 6 |
| 2024 | CoAst: Validation-Free Contribution Assessment for Federated Learning based on Cross-Round ValuationabstractIn the federated learning (FL) process, since the data held by each participant is different, it is necessary to figure out which participant has a higher contribution to the model performance. Effective contribution assessment can help motivate data owners to participate in the FL training. Research works in this field can be divided into two directions based on whether a validation dataset is required. Validation-based methods need to use representative validation data to measure the model accuracy, which is difficult to obtain in practical FL scenarios. Existing validation-free methods assess the contribution based on the parameters and gradients of local models and the global model in a single training round, which is easily compromised by the stochasticity of model training. In this work, we propose CoAst, a practical method to assess the FL participants' contribution without access to any validation data. The core idea of CoAst involves two aspects: one is to only count the most important part of model parameters through a weights quantization, and the other is a cross-round valuation based on the similarity between the current local parameters and the global parameter updates in several subsequent communication rounds. Extensive experiments show that CoAst has comparable assessment reliability to existing validation-based methods and outperforms existing validation-free methods. Hao Wu 0067, Shucheng Li, Fengyuan Xu, Sheng Zhong 0002 |
ACM Multimedia | 4 |
| 2024 | TTFL: Towards Trustworthy Federated Learning with Arm Confidential ComputingabstractFederated learning (FL), as a distributed training paradigm, has drawn great attention from both academia and industry. Recently, privacy and security concerns have been raised for FL. Despite many efforts to protect privacy and security, an FL framework that can systematically provide privacy and security guarantees is lacking. In this work, we present TTFL, a trustworthy FL framework in practice to defend the security and privacy issues based on Arm Confidential Compute Architecture (CCA). TTFL has two core designs. (1) It achieves a high-availability privacy protection based on flexible Trusted Execution Environments (TEEs). It leverages the resource-rich and conveniently accessed features of the latest TEE on Arm CCA, combined with our TEE secure interconnection design, to enable the whole FL process performed in distributed TEEs, which efficiently protects parameter confidentiality and protocol integrity. (2) It achieves effective security protection by proposing an effective poisoning-resisted secure aggregation scheme and protecting it within TEE. The new proposed secure aggregation combines the advantages of existing defenses and is placed in the flexible TEE to ensure a secure, effective, and non-bypassable aggregation procedure. We implement a prototype of TTFL and evaluate it regarding security, privacy, and system performance. Evaluation results show that TTFL can comprehensively and efficiently address the main privacy and security threats in FL. For instance, compared with previous work, it improves the model accuracy by 1.9% and reduces the attack success rate by 79.7% on the CIFAR-10 dataset with only about 19.8% training time overhead. Lizhi Sun, Jingzhou Zhu, Boyu Chang, Yixin Xu 0003, Hao Wu 0067, Fengyuan Xu, Sheng Zhong 0002 |
TrustCom | 7 |
| 2024 | TIM: Enabling Large-Scale White-Box Testing on In-App Deep Learning ModelsabstractIntelligent Applications (iApps), equipped with in-App deep learning (DL) models, are emerging to provide reliable DL inference services. However, in-App DL models are typically compiled into inference-only versions to enhance system performance, thereby impeding the evaluation of DL models. Specifically, the assessment of in-App models currently relies on black-box testing methods rather than direct white-box testing approaches. In this work, we propose TIM, an automated tool designed for conducting large-scale white-box testing of in-App models. Taking an iApp as input, TIM can lift the black-box (i.e., inference-only) in-App DL model into a backpropagation-enabled one and package it together, allowing comprehensive DL model testing or security issues detection. TIM proposes two reconstruction techniques to convert the inference-only model to a backpropagation-enabled version and reconstruct the DL-related IO processing code. In our experiments, we utilize TIM to extract 100 unique commercial in-App models and convert the models to white-box models, enabling backpropagation functionality. Experimental results show that TIM’s reconstruction techniques exhibit high accuracy. We open-source our prototype and part of the experimental data on the websitehttps://zenodo.org/record/7548141. Hao Wu 0067, Yuhang Gong, Xiaopeng Ke, Hanzhong Liang, Fengyuan Xu, Yunxin Liu 0001, Sheng Zhong 0002 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2024 | Multi-Label and Evolvable Dataset Preparation for Web-Based Object DetectionabstractIn this article, we focus on the emerging field of web-based object detection, which has gained considerable attention due to its ability to utilize large amounts of web data for training, thus eliminating the need for labor-intensive manual annotations. However, the noisy and ever-evolving nature of web data poses challenges in preparing high-quality datasets for web-based object detection. To address these challenges, we propose a fully automatic dataset preparation method in this article. Our proposed method incorporates a hierarchical clustering module that assigns multiple precise labels to each image. This module is based on our observation that web image data exhibits different distributions at varying granularities. Furthermore, an evolutionary relabeling module ensures the adaptability of both the prepared dataset and trained detection models to the ever-evolving web data. Extensive experiments demonstrate that our method outperforms other web-based methods, and achieves a comparable performance to those manually labeled benchmark datasets. Shucheng Li, Jingzhou Zhu, Boyu Chang, Hao Wu 0067, Fengyuan Xu, Sheng Zhong 0002 |
ACM Trans. Knowl. Discov. Data | 5 |
| 2024 | RescQR: Enabling Reliable Data Recovery in Screen-Camera Communication SystemabstractWith an increasing number of mobile devices equipped with screens and cameras, screen-camera communication (SCC) systems enable data exchange between devices conveniently and efficiently. By encoding data with spatial and temporal diversity on a screen, multiple users with a camera can receive data without setting up a wireless network. However, as the transmitter pushes the limits of increasing throughput with a high display rate, the receiver actually suffers from a low goodput caused by composite frames. Those frames cannot be decoded correctly with existing methods. To address this problem, we propose a reliable data recovery scheme named RescQR. In RescQR, a mixture separation scheme coupled with a dedicated frame border is proposed to separate composite frames. A Viterbi-based data recovery scheme is proposed to recover data from blurred regions in composite frames. Additionally, an auto-configuration method with the help of a front camera is proposed to adjust parameters automatically according to the estimated distance between the screen and the camera. Our prototype and experiments demonstrate that RescQR achieves a data goodput of 400+kbps even with standard QR codes, which significantly outperforms previous solutions. Kunming Xie, Xiaojun Zhu 0001, Yanchao Zhao, Fengyuan Xu |
IEEE Trans. Mob. Comput. | 6 |
| 2023 | Manipulating Transfer Learning for Property InferenceabstractTransfer learning is a popular method for tuning pretrained (upstream) models for different downstream tasks using limited data and computational resources. We study how an adversary with control over an upstream model used in transfer learning can conduct property inference attacks on a victim's tuned downstream model. For example, to infer the presence of images of a specific individual in the downstream training set. We demonstrate attacks in which an adversary can manipulate the upstream model to conduct highly effective and specific property inference attacks (AUC score > 0.9), without incurring significant performance loss on the main task. The main idea of the manipulation is to make the upstream model generate activations (intermediate features) with different distributions for samples with and without a target property, thus enabling the adversary to distinguish easily between downstream models trained with and without training examples that have the target property. Our code is available at https://github.com/yulongt23/Transfer-Inference. Yulong Tian, Fnu Suya, Anshuman Suri, Fengyuan Xu, David Evans 0001 |
CVPR | 4 |
| 2023 | GAPter: Gray-Box Data Protector for Deep Learning Inference Services at User SideabstractThe widespread deployment of Deep Learning Inference Services (DLISes) has raised people’s concerns about their data privacy being breached. Although data privacy enhancement has recently attracted a lot of attention, existing solutions all require the cooperation of service providers. Users lose control of their data when making data privacy enhancement decisions. However, it is difficult to enable the user-side control of data abuse prevention because users do not have any programming skills, deep learning knowledge, or rich computing resources. In this work, we propose a fully-automatic userside data privacy enhancement solution, GAPter, for DLISes. Given such a DLIS, GAPter can adaptively fuzz the service for a suitable enhancement strategy, with no cooperation between the DLIS provider and the user. We have implemented and comprehensively evaluated GAPter. The experimental results show that GAPter can find good balance points between privacy enhancement and user data utility. Hao Wu 0067, Xiaopeng Ke, Siyi He, Fengyuan Xu, Sheng Zhong 0002 |
ICASSP | 5 |
| 2023 | SAPPX: Securing COTS Binaries with Automatic Program Partitioning for Intel SGXabstractIn the era of cloud computing, many applications are migrated to public servers not fully controlled by users who may fear their critical operations or data from being compromised by attackers. Previous studies have shown that Intel SGX enclaves can improve applications’ security in many market products. Yet they mainly rely on developers to reprogram and recompile the application into an SGX-aware version. To address this problem, we propose SAPPX, an SGX-based program retrofitting method that can automatically partition COTS application binaries into two parts without breaking the original program semantics. The first part of the application runs in user space, while the second part is executed in an SGX enclave to protect the user’s sensitive information. We have implemented a prototype of SAPPX on x86/Linux platforms and evaluated its performance using real-world applications and SPECCPU 2017 benchmarks. The experimental results show that the average overhead of the proposed approach is up to 19%. Fengyuan Xu, Bing Chen 0002 |
ISSRE | 3 |
| 2023 | SIEGE: Self-Supervised Incremental Deep Graph Learning for Ethereum Phishing Scam DetectionabstractThe phishing scams pose a serious threat to the ecosystem of Ethereum which is one of the largest blockchains in the world. Such a type of cyberattack recently has caused losses of millions of dollars. In this paper, we propose a Self-supervised IncrEmental deep Graph lEarning (SIEGE) model, for the phishing scam detection problem on Ethereum. To overcome the data scalability challenge, we propose splitting the original Ethereum transaction data and constructing transaction graphs for each split. Confronted with the minimal labeled data available, we resort to graph-based self-supervised learning. We design a spatial pretext task to learn high-quality node embeddings inside a single graph split, as well as an incremental learning paradigm and a temporal pretext task to facilitate information flow between different graph splits. To evaluate the effectiveness of SIEGE, we gather a real-world dataset consisting of six-month Ethereum transaction records. The results demonstrate that our model consistently outperforms baseline approaches in both transductive and inductive settings. Shucheng Li, Runchuan Wang, Hao Wu 0067, Sheng Zhong 0002, Fengyuan Xu |
ACM Multimedia | 5 |
| 2023 | ShuffleCAN: Enabling Moving Target Defense for Attack Mitigation on Automotive CANabstractController Area Networks (CANs), the most widely used protocols for in-vehicle networks, are vulnerable to various attacks due to the lack of security countermeasures by design. CAN messages are broadcast without source/destination labeling and lack built-in encryption or authentication mechanisms, thus suffering many attacks. To address this problem, we propose a lightweight CAN message obfuscation technique called ShuffleCAN. Motivated by the idea of moving target defense (MTD), ShuffleCAN is designed with a combined shuffling scheme based on the hash chain and combinatorial coding techniques to achieve both ID anonymization and payload shuffling. With ShuffleCAN, selected or all transmitter and receiver pairs can communicate in a private dialect over the standard CAN protocol, so the eavesdropper cannot understand the meaning of each message or inject a valid fake message. We implemented a prototype and evaluated ShuffleCAN on Toyota’s testbed PASTA. The experimental results show that ShuffleCAN outperforms state-of-the-art CAN protection schemes. Huiping Qian, Xiaojun Zhu 0001, Fengyuan Xu |
MSN | 4 |
| 2023 | Dataset Preparation for Arbitrary Object Detection: An Automatic Approach based on Web Information in EnglishabstractAutomatic dataset preparation can help users avoid labor-intensive and costly manual data annotations. The difficulty in preparing a high-quality dataset for object detection involves three key aspects: relevance, naturality, and balance, which are not addressed by existing works. In this paper, we leverage information from the web, and propose a fully-automatic dataset preparation mechanism without any human annotation, which can automatically prepare a high-quality training dataset for the detection task with English text terms describing target objects. It contains three key designs, i.e., keyword expansion, data de-noising, and data balancing. Our experiments demonstrate that the object detectors trained with auto-prepared data are comparable to those trained with benchmark datasets and outperform other baselines. We also demonstrate the effectiveness of our approach in several more challenging real-world object categories that are not included in the benchmark datasets. Shucheng Li, Boyu Chang, Hao Wu 0067, Sheng Zhong 0002, Fengyuan Xu |
SIGIR | 6 |
| 2023 | RF-Badge: Vital Sign-Based Authentication via RFID Tag Array on BadgesabstractNowadays, authentication systems are usually required to provide continuous, contactless, and non-intrusive services. In this paper, we proposeRF-Badge, a vital sign-based authentication scheme on human subjects to meet the above requirements by using RFID technology. We consider two biometric features with individual diversity to characterize the vital sign of users, including themovement effectfrom respiration and thereflection effectfrom organs, especially the heart. To derive the movement effect from respiration, we build a phase-based geometric model to restore the fine-grained badge moving trace as the feature. To derive the reflection effect from human internal organs, we extract the reflection signal from the original signal and generate the spectrum as the feature. Besides, to deal with the feature deviation in different physical conditions of users, we propose a multi-condition network (MCNet) to further guarantee the generalization of RF-Badge. We implement a prototype system and evaluate the performance in real environments. The experiment results show that our system achieves the average false positive rate (FPR) of 3.9 percent and false negative rate (FNR) of 3.3 percent for continuous authentication within four signal cycles. Jingyi Ning, Lei Xie 0004, Yanling Bu, Fengyuan Xu, Da-Wei Zhou 0001, Sanglu Lu |
IEEE Trans. Mob. Comput. | 5 |
| 2023 | LEAP: TrustZone Based Developer-Friendly TEE for Intelligent Mobile AppsabstractARM TrustZone is widely deployed on commercial-off-the-shelf mobile devices for secure execution. However, many Apps cannot enjoy this feature because it brings many constraints to App developers. Previous works have been proposed to build a secure execution environment for developers on top of TrustZone. Unfortunately, these works are still not a fully-fledged solution for mobile Apps, especially for the emerging intelligent Apps. To this end, we propose LEAP, which is a lightweight developer-friendly TEE solution for mobile Apps. LEAP enables isolated codes to execute in parallel and access peripheral (e.g., mobile GPUs) with ease, flexibly manages system resources upon different workloads, and offers the auto DevOps tool to help developers prepare the codes running on it. We implement the LEAP prototype on the off-the-shelf ARM platform and conduct extensive experiments on it. The experimental results show that Apps can be adapted to run with LEAP easily and efficiently. Compared to the state-of-the-art work along this research line, LEAP can achieve an average 3.57× speedup in supporting intelligent Apps using mobile GPU acceleration. Lizhi Sun, Shuocheng Wang, Hao Wu 0067, Yuhang Gong, Fengyuan Xu, Yunxin Liu 0001, Sheng Zhong 0002 |
IEEE Trans. Mob. Comput. | 5 |
| 2023 | Privacy-Preserving Data Integrity Verification for Secure Mobile Edge StorageabstractMobile edge computing (MEC) is proposed as an extension of cloud computing in the scenarios where the end devices desire better services in terms of response time. Because the edges are usually owned by individuals or small organizations with limited operation capabilities, the data on the edges are easily corrupted (due to external attacks or internal hardware failures). Therefore, it is essential to verify data integrity in the MEC. We propose two Integrity Checking protocols for the mobile Edge storage, called ICE-basic and ICE-batch. Our protocols allow a third-party verifier to check the data integrity on the edges without violating users data privacy and query pattern privacy. We rigorously prove the security and privacy guarantees of the protocols. In addition, we have investigated how to let the end devices cache some verification tags such that the communication cost between end devices and the cloud can be further reduced when a user connects to multiple edges in sequence. We have implemented a proof-of-concept system that runs ICE, and extensive experiments are conducted to evaluate the performance of the proposed protocols. The theoretical analysis and experimental results demonstrate the proposed protocols are efficient both in computation and communication. Bingbing Jiang 0002, Fengyuan Xu, Qun Li 0001, Sheng Zhong 0002 |
IEEE Trans. Mob. Comput. | 4 |
| 2022 | Towards Automated Safety Vetting of Smart Contracts in Decentralized ApplicationsabstractWe propose VetSC, a novel UI-driven, program analysis guided model checking technique that can automatically extract contract semantics in DApps so as to enable targeted safety vetting. To facilitate model checking, we extract business model graphs from contract code that capture its intrinsic business and safety logic. To automatically determine what safety specifications to check, we retrieve textual semantics from DApp user interfaces. To exclude untrusted UI text, we also validate the UI-logic consistency and detect any discrepancies. We have implemented VetSC and applied it to 34 real-world DApps. Experiments have demonstrated that VetSC can accurately interpret smart contract code, enable autonomous safety vetting, and discover safety risks in real-world Dapps. Using our tool, we have successfully discovered 19 new safety risks in the wild, such as expired lottery tickets and double voting. Yue Duan, Shucheng Li, Minghao Li 0003, Fengyuan Xu, Mu Zhang 0001 |
CCS | 6 |
| 2022 | Towards Practical and Efficient Long Video SummaryabstractRecently, video summarization (VS) techniques are widely used to alleviate huge processing pressure brought by numerous long videos. However, it is hard to summarize long videos efficiently since processing hundreds of frames is still time-consuming. In this paper, we find that the Kernel Temporal Segmentation (KTS) method designed for detecting the shot boundaries in SOTA VS methods is time-consuming while handling long videos. To address this issue, we propose the Distribution-based KTS (D-KTS) by fully considering the characteristic of shot length distribution. Furthermore, we propose the Hash-based Adaptive Frame Selection (HAFS) to improve the system performance by fully taking advantage of the temporal locality of long videos. Our experiments present that the proposed D-KTS is 92.70% faster and takes up 90.08% less memory than the baseline KTS method on average. Xiaopeng Ke, Boyu Chang, Hao Wu 0067, Fengyuan Xu, Sheng Zhong 0002 |
ICASSP | 4 |
| 2022 | Privacy-Preserving and Robust Federated Deep Metric LearningabstractFederated learning, in contrast to traditional learning paradigms, has demonstrated its unique advantages in providing intelligence at the edge. However, existing federated learning approaches focus on the end-to-end classification tasks requiring a simple collaboration procedure where each participant can perform its local training independently. Unfortunately, there are still many tasks relying on learning the distinguishable feature metrics with respect to all the data, which is a different collaboration procedure across training participants. For example, the model for people identification has to ensure the feature representing a person is dissimilar to those representing others. To enable such federated learning for deep metrics (a.k.a federated deep metric learning) is challenging due to the data privacy and procedure robustness issues. With the consideration of these two challenges, this work proposes a novel computing framework for federated deep metric learning. This framework leverages the system-algorithm co-design to address privacy concerns via the Trusted Execution Environment (SGX enclave) and Differential Privacy mechanism. It also introduces a large-scale federated protocol which can robustly and efficiently deal with practical factors like the network fluctuation. We implement and evaluate our computing framework with two settings. One is a real-world implementation with a large number of mobile devices, while the other one is in our controllable environment for conducting experiments in various tasks. Our evaluation results show that our computing framework is able to train federated deep metric learning models with excellent scalability, data privacy preserving, and considerable accuracy even in exception conditions. Yulong Tian, Xiaopeng Ke, Zeyi Tao, Shaohua Ding, Fengyuan Xu, Qun Li 0001, Sheng Zhong 0002 |
IWQoS | 5 |
| 2022 | DIComP: Lightweight Data-Driven Inference of Binary Compiler Provenance with High AccuracyabstractBinary analysis is pervasively utilized to assess software security and test vulnerabilities without accessing source codes. The analysis validity is heavily influenced by the inferring ability of information related to the code compilation. Among the compilation information, compiler type and optimization level, as the key factors determining how binaries look like, are still difficult to be inferred efficiently with existing tools. In this paper, we conduct a thorough empirical study on the binary's appearance under various compilation settings and propose a lightweight binary analysis tool based on the simplest machine learning method, called DIComP to infer the compiler and optimization level via most relevant features according to the observation. Our comprehensive evaluations demonstrate that DIComP can fully recognize the compiler provenance, and it is effective in inferring the optimization levels with up to 90% accuracy. Also, it is efficient to infer thousands of binaries at a millisecond level with our lightweight machine learning model (1MB). Ligeng Chen, Zhongling He, Hao Wu 0067, Fengyuan Xu, Bing Mao 0001 |
SANER | 4 |
| 2022 | vTrust: Remotely Executing Mobile Apps Transparently With Local Untrusted OSabstractIncreasingly, many security and privacy sensitive applications (apps for short) are running in the mobile platforms. However, as the mobile operating systems are becoming increasingly sophisticated, they are vulnerable to various attacks. In addressing the need of running high assurance mobile apps in a secure environment even though the operating systems are untrusted, this paper presents VTRUST, a new mobile app trusted execution environment, which offloads the general execution and storage of a mobile app to a trusted remote server (e.g., a VM running in a cloud) and secures the I/O between the server and the mobile device with the aid of a trusted hypervisor on the mobile device. Specifically, VTRUST establishes an encrypted I/O channel between the local hypervisor and the remote server, such that any sensitive data flowing through the mobile OS, which is hosted by the hypervisor, is encrypted from the perspective of the local mobile OS. To enhance the performance of VTRUST, we have also designed multiple optimizations, such as output data compression and selective sensor data transmission. We have implemented VTRUST and our evaluation shows that it has limited impact on both user experience and the app performance. Yutao Tang, Zhengrui Qin, Zhiqiang Lin 0001, Yue Li 0002, Shanhe Yi, Fengyuan Xu, Qun Li 0001 |
IEEE Trans. Computers | 6 |
| 2022 | Stealthy Backdoors as Compression ArtifactsabstractModel compression is a widely-used approach for reducing the size of deep learning models without much accuracy loss, enabling resource-hungry models to be compressed for use on resource-constrained devices. In this paper, we study the risk that model compression could provide an opportunity for adversaries to inject stealthy backdoors. In a backdoor attack on a machine learning model, an adversary produces a model that performs well on normal inputs but outputs targeted misclassifications on inputs containing a small trigger pattern. We design stealthy backdoor attacks such that the full-sized model released by adversaries appears to be free from backdoors (even when tested using state-of-the-art techniques), but when the model is compressed it exhibits a highly effective backdoor. We show this can be done for two common model compression techniques—model pruning and model quantization—even in settings where the adversary has limited knowledge of how the particular compression will be done. Our findings demonstrate the importance of performing security tests on the models that will actually be deployed not in their precompressed version. Our implementation is available athttps://github.com/yulongtzzz/Stealthy-Backdoors-as-Compression-Artifacts. Yulong Tian, Fnu Suya, Fengyuan Xu, David Evans 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2021 | A System for Efficiently Hunting for Cyber Threats in Computer Systems Using Threat IntelligenceabstractLog-based cyber threat hunting has emerged as an important solution to counter sophisticated cyber attacks. However, existing approaches require non-trivial efforts of manual query construction and have overlooked the rich external knowledge about threat behaviors provided by open-source Cyber Threat Intelligence (OSCTI). To bridge the gap, we build ThreatRaptor, a system that facilitates cyber threat hunting in computer systems using OSCTI. Built upon mature system auditing frameworks, ThreatRaptor provides (1) an unsupervised, light-weight, and accurate NLP pipeline that extracts structured threat behaviors from unstructured OSCTI text, (2) a concise and expressive domain-specific query language, TBQL, to hunt for malicious system activities, (3) a query synthesis mechanism that automatically synthesizes a TBQL query from the extracted threat behaviors, and (4) an efficient query execution engine to search the big system audit logging data. Peng Gao 0008, Fei Shao, Xusheng Xiao, Fengyuan Xu, Prateek Mittal, Sanjeev R. Kulkarni, Dawn Song |
ICDE | 7 |
| 2021 | Enabling Efficient Cyber Threat Hunting With Cyber Threat IntelligenceabstractLog-based cyber threat hunting has emerged as an important solution to counter sophisticated attacks. However, existing approaches require non-trivial efforts of manual query construction and have overlooked the rich external threat knowledge provided by open-source Cyber Threat Intelligence (OSCTI). To bridge the gap, we propose ThreatRaptor, a system that facilitates threat hunting in computer systems using OSCTI. Built upon system auditing frameworks, ThreatRaptor provides (1) an unsupervised, light-weight, and accurate NLP pipeline that extracts structured threat behaviors from unstructured OSCTI text, (2) a concise and expressive domain-specific query language, TBQL, to hunt for malicious system activities, (3) a query synthesis mechanism that automatically synthesizes a TBQL query for hunting, and (4) an efficient query execution engine to search the big audit logging data. Evaluations on a broad set of attack cases demonstrate the accuracy and efficiency of ThreatRaptor in practical threat hunting. Peng Gao 0008, Fei Shao, Xusheng Xiao, Fengyuan Xu, Prateek Mittal, Sanjeev R. Kulkarni, Dawn Song |
ICDE | 6 |
| 2021 | Privacy-Preserving Optimal Recovering for the Nearly Exhausted Payment ChannelsabstractPayment Channel Network (PCN) is one of the most promising technologies for scaling the capacity of blockchain-based cryptocurrencies and improving the quality of blockchain-based services. However, during the use of PCNs, a significant portion of the payment channels gradually become exhausted, which triggers additional consumption of on-chain resources and makes PCNs less useful. This is a fundamental problem for blockchain-based cryptocurrencies, worthy of a thorough investigation.In this paper, we propose OPRE, a protocol for OPtimal off-chain REcovering of payment channels, to solve this problem. It is optimal in that it recovers the maximum number of nearly exhausted channels in the PCN. Furthermore, we consider users’ privacy concerns and design a privacy-preserving version of this protocol, so that users’ balance information does not need to be revealed. This protocol maintains optimality in recovering payment channels while providing cryptographically strong privacy guarantee. In addition to the theoretical design and analysis, we also implement OPRE and experimentally evaluate its performance. The results show that the OPRE protocol is both efficient and effective. Minze Xu, Yuan Zhang 0004, Fengyuan Xu, Sheng Zhong 0002 |
IWQoS | 3 |
| 2021 | AsyMo: scalable and efficient deep-learning inference on asymmetric mobile CPUsabstractOn-device deep learning (DL) inference has attracted vast interest. Mobile CPUs are the most common hardware for on-device inference and many inference frameworks have been developed for them. Yet, due to the hardware complexity, DL inference on mobile CPUs suffers from two common issues: the poor performance scalability on the asymmetric multiprocessor, and energy inefficiency. Manni Wang, Shaohua Ding, Ting Cao 0003, Yunxin Liu 0001, Fengyuan Xu |
MobiCom | 5 |
| 2021 | PECAM: privacy-enhanced video streaming and analytics via securely-reversible transformationabstractAs Video Streaming and Analytics (VSA) systems become increasingly popular, serious privacy concerns have risen on exposing too much unnecessary private information to the VSA providers. Yet, it is challenging to protect privacy while still preserving desired VSA features, i.e., effective analytics, forensic support, resource efficiency, and real-time execution. In this paper, we present a VSA privacy enhancement system (PECAM), which addresses the above challenge with no change in the VSA back-end. PECAM leverages a novel Generative Adversarial Network to perform the privacy-enhanced securely-reversible video transformation. PECAM also incorporates a couple of system optimizations into its VSA workflow to reduce network bandwidth usage and enable real-time processing on cameras. We implement our PECAM prototype on commodity hardware and evaluate its performance via both security study and extensive experiments. Results demonstrate that PECAM can effectively enhance the visual privacy of VSA in the presence of an adversary, and its transformed videos, when taken as input for various VSA back-end tasks, maintain a 96% accuracy of corresponding original videos. Additionally, it performs 12.3× and 1.8× better than baseline methods in terms of the computing cost and network bandwidth usage, respectively. Hao Wu 0067, Xuejin Tian, Minghao Li 0003, Yunxin Liu 0001, Ganesh Ananthanarayanan, Fengyuan Xu, Sheng Zhong 0002 |
MobiCom | 6 |
| 2021 | DAPter: Preventing User Data Abuse in Deep Learning Inference ServicesabstractThe data abuse issue has risen along with the widespread development of the deep learning inference service (DLIS). Specifically, mobile users worry about their input data being labeled to secretly train new deep learning models that are unrelated to the DLIS they subscribe to. This unique issue, unlike the privacy problem, is about the rights of data owners in the context of deep learning. However, preventing data abuse is demanding when considering the usability and generality in the mobile scenario. In this work, we propose, to our best knowledge, the first data abuse prevention mechanism called DAPter. DAPter is a user-side DLIS-input converter, which removes unnecessary information with respect to the targeted DLIS. The converted input data by DAPter maintains good inference accuracy and is difficult to be labeled manually or automatically for the new model training. DAPter’s conversion is empowered by our lightweight generative model trained with a novel loss function to minimize abusable information in the input data. Furthermore, adapting DAPter requires no change in the existing DLIS backend and models. We conduct comprehensive experiments with our DAPter prototype on mobile devices and demonstrate that DAPter can substantially raise the bar of the data abuse difficulty with little impact on the service quality and overhead. Hao Wu 0067, Xuejin Tian, Yuhang Gong, Minghao Li 0003, Fengyuan Xu |
WWW | 6 |
| 2021 | Towards Thwarting Template Side-Channel Attacks in Secure Cloud DeduplicationsabstractAs one of a few critical technologies to cloud storage service, deduplication allows cloud servers to save storage space by deleting redundant file copies. However, it often leaks side channel information regarding whether an uploading file gets deduplicated or not. Exploiting this information, adversaries can easily launch a template side-channel attack and severely harm cloud users' privacy. To thwart this kind of attack, we resort to the k-anonymity privacy concept to design secure threshold deduplication protocols. Specifically, we have devised a novel cryptographic primitive called “dispersed convergent encryption” (DCE) scheme, and proposed two different constructions of it. With these DCE schemes, we successfully construct secure threshold deduplication protocols that do not rely on any trusted third party. Our protocols not only support confidentiality protections and ownership verifications, but also enjoy formal security guarantee against template side-channel attacks even when the cloud server could be a “covert adversary” who may violate the predefined threshold and perform deduplication covertly. Experimental evaluations show our protocols enjoy very good performance in practice. Yuan Zhang 0004, Yunlong Mao, Minze Xu, Fengyuan Xu, Sheng Zhong 0002 |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2020 | Detecting GAN-based Privacy Attack in Distributed LearningabstractDistributed learning unleashes the power of training collaboration among multiple parties who have different training data. While participants enjoy mutually beneficial outcomes of the distributed learning, which cannot be achieved by single party, they also worry about the risk of privacy leaking. In fact, recent work shows that a malicious participant is able to leverage a Generative Adversarial Network (GAN) to steal sensitive information of training data owned by others through shared gradient updates. However, existed countermeasures, such as the differential privacy or cryptographic methods, could disturb the training in terms of model accuracy or computation overhead. In this paper, we seek to mitigate this privacy issue in a non-intrusive manner. Instead of passive protection, we propose to actively detect such GAN-based attackers at the very beginning of training. Our detection only utilizes the gradient updates uploaded by participants during the training, so it is transparent to participants and does not require protocol changes. We demonstrate the effectiveness of our detection through extensive experiments in different settings and attack scenarios. Yayuan Xiong, Fengyuan Xu, Sheng Zhong 0002 |
ICC | 2 |
| 2020 | Exploiting Adversarial Examples to Drain Computational Resources on Mobile Deep Learning SystemsabstractIn order to perform deep learning tasks everywhere, many optimizations have been proposed to address the resource limitations on mobile systems like IoTs. A key approach among others is to dynamically adjust computational resources of the deep learning inference according to the characteristics of incoming inputs. For example, one of popular optimizations is to pick for each input a suitable combination of computations with respect to its inference difficulty. However, we find out that such “dynamic routing” of computations could be exploited to drain/waste precious resources on mobile deep learning systems. In this work, we introduce a new deep learning attack dimension, the computational resources draining, and demonstrate its feasibility in one of possible attack manners, the adversarial examples of input data. We describe how to construct our special adversarial examples aiming to the resource draining, and show that these poisoned inputs are able to increase the computation loads on purpose with two experiment datasets. We hope that our findings can shed light on the path of improving the robustness of mobile deep learning optimizations. Yulong Tian, Rongchun Yao, Fengyuan Xu, Sheng Zhong 0002 |
SEC | 4 |
| 2020 | Efficient Architecture Paradigm for Deep Learning Inference as a ServiceabstractDeep learning (DL) inference has been broadly used and shown excellent performance in many intelligent applications. Unfortunately, the high resource consumption and training efforts of sophisticated models present obstacles for regular users to enjoy it. Thus, Deep Learning Inference as a Service (DIaaS), offering online inference services on cloud, has earned great popularity among cloud tenants who can send their DIaaS inputs via RPCs across the internal network. However, such detached architecture paradigm is inappropriate to DIaaS because the high-dimensional inputs of DIaaS consume a lot of precious internal bandwidth and the service latency of DIaaS has to be low and stable. We therefore propose a novel architecture paradigm on cloud for DIaaS in order to address the above two problems without giving up the security and maintenance benefits. We first leverage the SGX technology, a strongly-protected user space enclave, to bring DIaaS computation to its input source as close as possible, i.e. co-locating a cloud tenant and its subscribed DIaaS in the same virtual machine. When the GPU acceleration is needed, we migrate this virtual machine to any available GPU host and transparently utilize the GPU via our backend computing stack installed on it. In this way the majority of internal bandwidth is saved compared to traditional paradigm. Furthermore, we greatly improve the efficiency of the proposed architecture paradigm, from the computation and I/O perspectives, by making the entire data flow more DL-oriented. Finally, we implement a prototype system and evaluate it in real-world scenarios. The experiments show that our locality-aware architecture achieves the average single CPU (GPU) based deep learning inference time 2.84X (4.87X) less than the traditional detached architecture on average. Xiaopeng Ke, Fengyuan Xu |
IPCCC | 3 |
| 2020 | EMO: real-time emotion recognition from single-eye images for resource-constrained eyewear devicesabstractReal-time user emotion recognition is highly desirable for many applications on eyewear devices like smart glasses. However, it is very challenging to enable this capability on such devices due to tightly constrained image contents (only eye-area images available from the on-device eye-tracking camera) and computing resources of the embedded system. In this paper, we propose and develop a novel system called EMO that can recognize, on top of a resource-limited eyewear device, real-time emotions of the user who wears it. Unlike most existing solutions that require whole-face images to recognize emotions, EMO only utilizes the single-eye-area images captured by the eye-tracking camera of the eyewear. To achieve this, we design a customized deep-learning network to effectively extract emotional features from input single-eye images and a personalized feature classifier to accurately identify a user's emotions. EMO also exploits the temporal locality and feature similarity among consecutive video frames of the eye-tracking camera to further reduce the recognition latency and system resource usage. We implement EMO on two hardware platforms and conduct comprehensive experimental evaluations. Our results demonstrate that EMO can continuously recognize seven-type emotions at 12.8 frames per second with a mean accuracy of 72.2%, significantly outperforming the state-of-the-art approach, and consume much fewer system resources. Hao Wu 0067, Xuejin Tian, Edward Sun, Yunxin Liu 0001, Fengyuan Xu, Sheng Zhong 0002 |
MobiSys | 7 |
| 2020 | Escaping Backdoor Attack Detection of Deep Learning
Yayuan Xiong, Fengyuan Xu, Sheng Zhong 0002, Qun Li 0001 |
SEC | 2 |
| 2019 | DeepIntent: Deep Icon-Behavior Learning for Detecting Intention-Behavior Discrepancy in Mobile AppsabstractMobile apps have been an indispensable part in our daily life. However, there exist many potentially harmful apps that may exploit users' privacy data, e.g., collecting the user's information or sending messages in the background. Keeping these undesired apps away from the market is an ongoing challenge. While existing work provides techniques to determine what apps do, e.g., leaking information, little work has been done to answer, are the apps' behaviors compatible with the intentions reflected by the app's UI? In this work, we explore the synergistic cooperation of deep learning and program analysis as the first step to address this challenge. Specifically, we focus on the UI widgets that respond to user interactions and examine whether the intentions reflected by their UIs justify their permission uses. We present DeepIntent, a framework that uses novel deep icon-behavior learning to learn an icon-behavior model from a large number of popular apps and detect intention-behavior discrepancies. In particular, DeepIntent provides program analysis techniques to associate the intentions (i.e., icons and contextual texts) with UI widgets' program behaviors, and infer the labels (i.e., permission uses) for the UI widgets based on the program behaviors, enabling the construction of a large-scale high-quality training dataset. Based on the results of the static analysis, DeepIntent uses deep learning techniques that jointly model icons and their contextual texts to learn an icon-behavior model, and detects intention-behavior discrepancies by computing the outlier scores based on the learned model. We evaluate DeepIntent on a large-scale dataset (9,891 benign apps and 16,262 malicious apps). With 80% of the benign apps for training and the remaining for evaluation, DeepIntent detects discrepancies with AUC scores 0.8656 and 0.8839 on benign apps and malicious apps, achieving 39.9% and 26.1% relative improvements over the state-of-the-art approaches. Shengqu Xi, Shao Yang, Xusheng Xiao, Yuan Yao 0001, Yayuan Xiong, Fengyuan Xu, Haoyu Wang 0001, Peng Gao 0008, Zhuotao Liu, Feng Xu 0007, Jian Lu 0001 |
CCS | 6 |
| 2019 | Privacy-Preserving Data Integrity Verification in Mobile Edge ComputingabstractMobile edge computing (MEC) is proposed as an extension of cloud computing in the scenarios where the end devices desire better services in terms of response time. Edge nodes are deployed at the proximity of the end devices, and it can pre-download parts of data stored in the cloud so that the end devices can access these data with low latency. However, because the edges are usually owned by individuals and small organizations, which have limited operation capacities for maintaining the machines, the data on the edges are easily corrupted (due to external attacks or internal hardware failures). Therefore, it is essential to verify data integrity in the MEC. We propose two Integrity Checking protocols for mobile Edge computing, called ICE-basic and ICE-batch, which are designed for the cases where the user wants to check data integrity on a single edge or multiple edges, respectively. Based on the concept of provable data possession and the technique of private information retrieval, our protocols allow a third-party verifier to check the data integrity on the edges without violating users' data privacy and query pattern privacy. We rigorously prove the security and privacy guarantees of the protocols. Furthermore, we have implemented a proof-of-concept system that runs ICE, and extensive experiments are conducted. The theoretical analysis and experimental results demonstrate the proposed protocols are efficient both in computation and communication. Bingbing Jiang 0002, Fengyuan Xu, Qun Li 0001, Sheng Zhong 0002 |
ICDCS | 3 |
| 2019 | Occlumency: Privacy-preserving Remote Deep-learning Inference Using SGXabstractDeep-learning (DL) is receiving huge attention as enabling techniques for emerging mobile and IoT applications. It is a common practice to conduct DNN model-based inference using cloud services due to their high computation and memory cost. However, such a cloud-offloaded inference raises serious privacy concerns. Malicious external attackers or untrustworthy internal administrators of clouds may leak highly sensitive and private data such as image, voice and textual data. In this paper, we propose Occlumency, a novel cloud-driven solution designed to protect user privacy without compromising the benefit of using powerful cloud resources. Occlumency leverages secure SGX enclave to preserve the confidentiality and the integrity of user data throughout the entire DL inference process. DL inference in SGX enclave, however, impose a severe performance degradation due to limited physical memory space and inefficient page swapping. We designed a suite of novel techniques to accelerate DL inference inside the enclave with a limited memory size and implemented Occlumency based on Caffe. Our experiment with various DNN models shows that Occlumency improves inference speed by 3.6x compared to the baseline DL inference in SGX and achieves a secure DL inference within 72% of latency overhead compared to inference in the native environment. Taegyeong Lee, Saumay Pushp, Caihua Li, Yunxin Liu 0001, Youngki Lee 0001, Fengyuan Xu, Chenren Xu, Junehwa Song |
MobiCom | 7 |
| 2019 | Trojan Attack on Deep Generative Models in Autonomous Driving
Shaohua Ding, Yulong Tian, Fengyuan Xu, Qun Li 0001, Sheng Zhong 0002 |
SecureComm (1) | 3 |
| 2019 | A Query System for Efficiently Investigating Complex Attack Behaviors for Enterprise SecurityabstractThe need for countering Advanced Persistent Threat (APT) attacks has led to the solutions that ubiquitously monitor system activities in each enterprise host, and perform timely attack investigation over the monitoring data for uncovering the attack sequence. However, existing general-purpose query systems lack explicit language constructs for expressing key properties of major attack behaviors, and their semantics-agnostic design often produces inefficient execution plans for queries. To address these limitations, we build Aiql, a novel query system that is designed with novel types of domain-specific optimizations to enable efficient attack investigation. Aiql provides (1) a domain-specific data model and storage for storing the massive system monitoring data, (2) a domain-specific query language, Attack Investigation Query Language (Aiql) that integrates critical primitives for expressing major attack behaviors, and (3) an optimized query engine based on the characteristics of the data and the semantics of the query to efficiently schedule the execution. We have deployed Aiql in NEC Labs America comprising 150 hosts. In our demo, we aim to show the complete usage scenario of Aiql by (1) performing an APT attack in a controlled environment, and (2) using Aiql to investigate such attack by querying the collected system monitoring data that contains the attack traces. The audience will have the option to perform the APT attack themselves under our guidance, and interact with the system and investigate the attack via issuing queries and checking the query results through our web UI. Peng Gao 0008, Xusheng Xiao, Zhichun Li, Kangkook Jee, Fengyuan Xu, Sanjeev R. Kulkarni, Prateek Mittal |
Proc. VLDB Endow. | 5 |
| 2018 | NodeMerge: Template Based Efficient Data Reduction For Big-Data Causality AnalysisabstractToday's enterprises are exposed to sophisticated attacks, such as Advanced Persistent Threats~(APT) attacks, which usually consist of stealthy multiple steps. To counter these attacks, enterprises often rely on causality analysis on the system activity data collected from a ubiquitous system monitoring to discover the initial penetration point, and from there identify previously unknown attack steps. However, one major challenge for causality analysis is that the ubiquitous system monitoring generates a colossal amount of data and hosting such a huge amount of data is prohibitively expensive. Thus, there is a strong demand for techniques that reduce the storage of data for causality analysis and yet preserve the quality of the causality analysis. To address this problem, in this paper, we propose NodeMerge, a template based data reduction system for online system event storage. Specifically, our approach can directly work on the stream of system dependency data and achieve data reduction on the read-only file events based on their access patterns. It can either reduce the storage cost or improve the performance of causality analysis under the same budget. Only with a reasonable amount of resource for online data reduction, it nearly completely preserves the accuracy for causality analysis. The reduced form of data can be used directly with little overhead. To evaluate our approach, we conducted a set of comprehensive evaluations, which show that for different categories of workloads, our system can reduce the storage capacity of raw system dependency data by as high as 75.7 times, and the storage capacity of the state-of-the-art approach by as high as 32.6 times. Furthermore, the results also demonstrate that our approach keeps all the causality analysis information and has a reasonably small overhead in memory and hard disk. Yutao Tang, Ding Li 0001, Zhichun Li, Mu Zhang 0001, Kangkook Jee, Xusheng Xiao, Zhenyu Wu 0003, Junghwan Rhee, Fengyuan Xu, Qun Li 0001 |
CCS | 9 |
| 2018 | MobiCrowd: Mobile Crowdsourcing on Location-based Social NetworksabstractThe great potential of mobile crowdsourcing has started to attract attention of both industries and the research community. However, current commercial mobile crowdsourcing marketplaces are unsatisfactory because of the limited worker base and functionality. In this paper, we first revisit the foundation of performing mobile crowdsourcing on location-based social networks (LBSNs) through specially designed survey studies and comparison experiments involving hundreds of users. Our results reveal that active check-ins are good indicators of picking a right user to perform tasks, and LBSN could be an ideal platform for mobile crowdsourcing given proper services provided. We then propose both the centralized and decentralized design of MobiCrowd, a mobile crowdsourcing service built on LBSNs. Our evaluation, through trace-driven simulation and real-world experiments, demonstrates that the proposed schemes can effectively find workers for mobile crowdsourcing tasks associated with different venues by analyzing their location check-in histories. Yulong Tian, Qun Li 0001, Fengyuan Xu, Sheng Zhong 0002 |
INFOCOM | 4 |
| 2018 | Internet Protocol Cameras with No Password Protection: An Empirical Investigation
Haitao Xu 0002, Fengyuan Xu, Bo Chen 0028 |
PAM | 2 |
| 2018 | AIQL: Enabling Efficient Attack Investigation from System Monitoring Data
Peng Gao 0008, Xusheng Xiao, Zhichun Li, Fengyuan Xu, Sanjeev R. Kulkarni, Prateek Mittal |
USENIX ATC | 4 |
| 2018 | Location privacy in public access points positioning: An optimization and geometry approach
Yunlong Mao, Yuan Zhang 0004, Fengyuan Xu, Sheng Zhong 0002 |
Comput. Secur. | 4 |
| 2018 | A Geo-Indistinguishable Location Perturbation Mechanism for Location-Based Services Supporting Frequent QueriesabstractAs location-based services (LBSs) on smartphones become increasingly popular, such services are causing serious privacy concerns, because many users are unwilling to see their location information leaked to service providers. Recently, in order to protect users’ location privacy, researchers have introducedgeo-indistinguishability, the first specialized privacy model for LBSs that can provide provable privacy guarantees. Intuitively, geo-indistinguishability means that through perturbation, any two locations within a given distance produce observations with similar distributions, and thus, attackers have no way to learn users’ real locations. However, even if geo-indistinguishability is achieved, there remains a significant threat to users’ location privacy: the privacy consumption increases with the number of queries for the existing geo-indistinguishable location perturbation mechanism, and therefore, there is a high risk of privacy violation when the number of queries is not small. In this paper, we enhance the privacy protection for LBSs by proposing an improved geo-indistinguishable mechanism. It can reduce the privacy costs to almost 0 when the user’s location satisfies a condition. We also present an improvement to further reduce the privacy costs when the above condition is not satisfied. Evaluations upon two public trace data sets show that the proposed mechanisms can dramatically save the privacy budget and thus support much more queries. The results also show that the proposed mechanisms are efficient, and their performance is controllable. Jingyu Hua, Fengyuan Xu, Sheng Zhong 0002 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2017 | Using Wireless Link Dynamics to Extract a Secret Key in Vehicular ScenariosabstractSecuring a wireless channel between any two vehicles is a crucial component of vehicular networks security. This can be done by using a secret key to encrypt the messages. We propose a scheme to allow two cars to extract a shared secret from RSSI (Received Signal Strength Indicator) values in such a way that nearby cars cannot obtain the same key. The key is information-theoretically secure, i.e., it is secure against an adversary with unlimited computing power. Although there are existing solutions of key extraction in the indoor or low-speed environments, the unique channel conditions make them inapplicable to vehicular environments. Our scheme effectively and efficiently handles the high noise and mismatch features of the measured samples so that it can be executed in the noisy vehicular environment. We also propose an online parameter learning mechanism to adapt to different channel conditions. Extensive real-world experiments are conducted to validate our solution. Xiaojun Zhu 0001, Fengyuan Xu, Edmund Novak, Chiu C. Tan 0001, Qun Li 0001, Guihai Chen |
IEEE Trans. Mob. Comput. | 2 |
| 2016 | High Fidelity Data Reduction for Big Data Security Dependency AnalysesabstractIntrusive multi-step attacks, such as Advanced Persistent Threat (APT) attacks, have plagued enterprises with significant financial losses and are the top reason for enterprises to increase their security budgets. Since these attacks are sophisticated and stealthy, they can remain undetected for years if individual steps are buried in background "noise." Thus, enterprises are seeking solutions to "connect the suspicious dots" across multiple activities. This requires ubiquitous system auditing for long periods of time, which in turn causes overwhelmingly large amount of system audit events. Given a limited system budget, how to efficiently handle ever-increasing system audit logs is a great challenge. This paper proposes a new approach that exploits the dependency among system events to reduce the number of log entries while still supporting high-quality forensic analysis. In particular, we first propose an aggregation algorithm that preserves the dependency of events during data reduction to ensure the high quality of forensic analysis. Then we propose an aggressive reduction algorithm and exploit domain knowledge for further data reduction. To validate the efficacy of our proposed approach, we conduct a comprehensive evaluation on real-world auditing systems using log traces of more than one month. Our evaluation results demonstrate that our approach can significantly reduce the size of system logs and improve the efficiency of forensic analysis without losing accuracy. Zhang Xu, Zhenyu Wu 0003, Zhichun Li, Kangkook Jee, Junghwan Rhee, Xusheng Xiao, Fengyuan Xu, Haining Wang 0001, Guofei Jiang |
CCS | 7 |
| 2013 | Fast Mencius: Mencius with low commit latencyabstractMencius is a protocol for general state machine replication that tolerates crash failures. It has high performance in wide-area networks. However, the commit latency of Mencius is limited by the slowest replica. This paper presents Fast Mencius, a crash fault-tolerant state machine replication protocol, which enhances Mencius with Active Revoke and Multi-instance Propose. Active Revoke allows the non-slow replicas to proceed without being delayed by the slowest replica, while Multi-instance Propose enables the slow replicas to have their proposals chosen by the replicated state machine. Our evaluation shows that in presence of slow replicas, Fast Mencius's commit latency is significantly lower than that of Mencius, and it also achieves high throughput. Harry Gao, Fengyuan Xu, Qun Li 0001 |
INFOCOM | 3 |
| 2013 | Extracting secret key from wireless link dynamics in vehicular environmentsabstractA crucial component of vehicular network security is to establish a secure wireless channel between any two vehicles. In this paper, we propose a scheme to allow two cars to extract a secret key from RSSI (Received Signal Strength Indicator) values in such a way that nearby cars cannot obtain the same secret. Our solution can be executed in noisy, outdoor vehicular environments. We also propose an online parameter learning mechanism to adapt to different channel conditions. We conduct extensive realworld experiments to validate our solution. Xiaojun Zhu 0001, Fengyuan Xu, Edmund Novak, Chiu C. Tan 0001, Qun Li 0001, Guihai Chen |
INFOCOM | 2 |
| 2013 | Optimizing background email sync on smartphonesabstractEmail is a key application used on smartphones. Even when the phone is in stand-by mode, users expect the phone to continue syncing with an email server to receive new mes-sages. Each such sync operation wakes up the smartphone for data reception and processing. In this paper, we show that this "cost of email sync" in stand-by mode constitutes a significant source of energy consumption, and thus reduces battery life. We quantify the power performance of different existing email clients on two smartphone platforms, An-droid and Windows Phone, and study the impact of system parameters such as email size, inbox size, and pull vs. push. Our results show that existing email clients do not handle email sync in an energy efficient way. This is because the underlying protocols and architectures are not designed for the specific needs of operating in stand-by mode. Based on our findings, we derive general design principles for energy-efficient event handling on smartphones, and apply these principles to the case of email sync and implement our techniques on commercial smartphones. Experimental results show that our techniques are able to significantly reduce energy cost of email sync by 49.9% on average with our experiment settings. Fengyuan Xu, Yunxin Liu 0001, Thomas Moscibroda, Ranveer Chandra, Yongguang Zhang, Qun Li 0001 |
MobiSys | 1 |
| 2013 | V-edge: Fast Self-constructive Power Modeling of Smartphones Based on Battery Voltage Dynamics
Fengyuan Xu, Yunxin Liu 0001, Qun Li 0001, Yongguang Zhang |
NSDI | 1 |
| 2013 | SybilDefender: A Defense Mechanism for Sybil Attacks in Large Social NetworksabstractDistributed systems without trusted identities are particularly vulnerable to sybil attacks, where an adversary creates multiple bogus identities to compromise the running of the system. This paper presents SybilDefender, a sybil defense mechanism that leverages the network topologies to defend against sybil attacks in social networks. Based on performing a limited number of random walks within the social graphs, SybilDefender is efficient and scalable to large social networks. Our experiments on two 3,000,000 node real-world social topologies show that SybilDefender outperforms the state of the art by more than 10 times in both accuracy and running time. SybilDefender can effectively identify the sybil nodes and detect the sybil community around a sybil node, even when the number of sybil nodes introduced by each attack edge is close to the theoretically detectable lower bound. Besides, we propose two approaches to limiting the number of attack edges in online social networks. The survey results of our Facebook application show that the assumption made by previous work that all the relationships in social networks are trusted does not apply to online social networks, and it is feasible to limit the number of attack edges in online social networks by relationship rating. Fengyuan Xu, Chiu C. Tan 0001, Qun Li 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2013 | SmartAssoc: Decentralized Access Point Selection Algorithm to Improve ThroughputabstractAs the first step of the communication procedure in 802.11, an unwise selection of the access point (AP) hurts one client's throughput. This performance downgrade is usually hard to be offset by other methods, such as efficient rate adaptations. In this paper, we study this AP selection problem in a decentralized manner, with the objective of maximizing the minimum throughput among all clients. We reveal through theoretical analysis that the selfish strategy, which commonly applies in decentralized systems, cannot effectively achieve this objective. Accordingly, we propose an online AP association strategy that not only achieves a minimum throughput (among all clients) that is provably close to the optimum, but also works effectively in practice with reasonable computation and transmission overhead. The association protocol applying this strategy is implemented on the commercial hardware and compatible with legacy APs without any modification. We demonstrate its feasibility and performance through real experiments and intensive simulations. Fengyuan Xu, Xiaojun Zhu 0001, Chiu C. Tan 0001, Qun Li 0001, Guanhua Yan, Jie Wu 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2012 | MobiShare: Flexible privacy-preserving location sharing in mobile online social networksabstractLocation sharing is a fundamental component of mobile online social networks (mOSNs), which also raises significant privacy concerns. The mOSNs collect a large amount of location information over time, and the users' location privacy is compromised if their location information is abused by adversaries controlling the mOSNs. In this paper, we present MobiShare, a system that provides flexible privacy-preserving location sharing in mOSNs. MobiShare is flexible to support a variety of location-based applications, in that it enables location sharing between both trusted social relations and untrusted strangers, and it supports range query and user-defined access control. In MobiShare, neither the social network server nor the location server has a complete knowledge of the users' identities and locations. The users' location privacy is protected even if either of the entities colludes with malicious users. Fengyuan Xu, Qun Li 0001 |
INFOCOM | 2 |
| 2012 | SybilDefender: Defend against sybil attacks in large social networksabstractDistributed systems without trusted identities are particularly vulnerable to sybil attacks, where an adversary creates multiple bogus identities to compromise the running of the system. This paper presents SybilDefender, a sybil defense mechanism that leverages the network topologies to defend against sybil attacks in social networks. Based on performing a limited number of random walks within the social graphs, SybilDefender is efficient and scalable to large social networks. Our experiments on two 3,000,000 node real-world social topologies show that SybilDefender outperforms the state of the art by one to two orders of magnitude in both accuracy and running time. SybilDefender can effectively identify the sybil nodes and detect the sybil community around a sybil node, even when the number of sybil nodes introduced by each attack edge is close to the theoretically detectable lower bound. Besides, we propose two approaches to limiting the number of attack edges in online social networks. The survey results of our Facebook application show that the assumption made by previous work that all the relationships in social networks are trusted does not apply to online social networks, and it is feasible to limit the number of attack edges in online social networks by relationship rating. Fengyuan Xu, Chiu C. Tan 0001, Qun Li 0001 |
INFOCOM | 2 |
| 2011 | Defending against vehicular rogue APsabstractThis paper considers vehicular rogue access points (APs) that rogue APs are set up in moving vehicles to mimic legitimate roadside APs to lure users to associate to them. Due to its mobility, a vehicular rogue AP is able to maintain a long connection with users. Thus, the adversary has more time to launch various attacks to steal users' private information. We propose a practical detection scheme based on the comparison of Receive Signal Strength (RSS) to prevent users from connecting to rogue APs. The basic idea of our solution is to force APs (both legitimate and fake) to report their GPS locations and transmission powers in beacons. Based on such information, users can validate whether the measured RSS matches the value estimated from the AP's location, transmission power, and its own GPS location. Furthermore, we consider the impact of path loss and shadowing and propose a method based on rate adaption to deal with advanced rogue APs. We implemented our detection technique on commercial off-the-shelf devices including wireless cards, antennas, and GPS modules to evaluate the efficacy of our scheme. Fengyuan Xu, Chiu C. Tan 0001, Yifan Zhang 0002, Qun Li 0001 |
INFOCOM | 2 |
| 2011 | IMDGuard: Securing implantable medical devices with the external wearable guardianabstractRecent studies have revealed security vulnerabilities in implantable medical devices (IMDs). Security design for IMDs is complicated by the requirement that IMDs remain operable in an emergency when appropriate security credentials may be unavailable. In this paper, we introduce IMDGuard, a comprehensive security scheme for heart-related IMDs to fulfill this requirement. IMDGuard incorporates two techniques tailored to provide desirable protections for IMDs. One is an ECG based key establishment without prior shared secrets, and the other is an access control mechanism resilient to adversary spoofing attacks. The security and performance of IMDGuard are evaluated on our prototype implementation. Fengyuan Xu, Zhengrui Qin, Chiu C. Tan 0001, Qun Li 0001 |
INFOCOM | 1 |
| 2010 | Designing a Practical Access Point Association ProtocolabstractIn a Wireless Local Area Network (WLAN), the Access Point (AP) selection of a client heavily influences the performance of its own and others. Through theoretical analysis, we reveal that previously proposed association protocols are not effective in maximizing the minimal throughput among all clients. Accordingly, we propose an online AP association strategy that not only achieves a minimal throughput (among all clients) that is provably close to the optimum, but also works effectively in practice with a reasonable computational overhead. The association protocol applying this strategy is implemented on the commercial hardware and compatible with legacy APs without any modification. We demonstrate its feasibility and performance through real experiments. Fengyuan Xu, Chiu C. Tan 0001, Qun Li 0001, Guanhua Yan, Jie Wu 0001 |
INFOCOM | 1 |
| 2009 | Experimental Study on Secure Data Collection in Vehicular Sensor Networks
Harry Gao, Seth Utecht, Fengyuan Xu, Qun Li 0001 |
WASA | 3 |