Xiangqun Chen

dblp:49/628 · DBLP profile ↗
← Back
82ranked-venue papers
0as first author
37since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 34 · 14 since 2021Applied, interdisciplinary, general and emerging computing · 20 · 5 since 2021Security and privacy · 16 · 10 since 2021Systems, architecture and hardware · 7 · 3 since 2021Computer networks · 7 · 4 since 2021Databases, data management, data science and information retrieval · 7 · 2 since 2021Artificial intelligence and machine learning · 4 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Bridging the Memory Hotness Gap in Edge Systems with Hotness-Segregated Object Allocation
abstract
Kernel operations in resource-constrained edge systems, such as memory swapping and deduplication, use the access frequency (hotness) of memory pages to guide page placement and reclamation. However, these operations suffer from page-hotness skew: a page may contain a mix of highly accessed and infrequently accessed objects, which causes inaccurate page-level classification, wasted DRAM capacity, and expensive I/O. We attribute this skewness to a cross-layer mismatch: the kernel manages memory at page granularity, whereas user-level allocators place objects without considering access hotness.
Ruizhe Huang, Jiahua Wang, Qihang Xu, Peng Jiang 0007, Zhida An, Ding Li 0001, Yao Guo 0001, Xiangqun Chen, Yuxin Ren 0001, Ning Jia 0004
LCTES8
2026 SymFlow: Event-Chain-Aware Symbolic Execution for Serverless Sensitive Data Flow Detection
abstract
Serverless applications are widely adopted for their scalability, cost-efficiency, and elastic resource management. However, their event-driven nature introduces complex event chains whose trigger-handler relationships are often determined dynamically by conditional logic, asynchronous callbacks, and resource-state dependencies. Existing security analysis tools, such as CloudFlow, mainly rely on static analysis, making it difficult to capture these dynamic event-chain interactions and the semantics of coarse-grained cloud APIs. As a result, they often fail to bridge the gap between architectural reachability and semantic feasibility, leading to both false positives and false negatives.
Yuanpeng Wang, Zhineng Zhong, Zhenkai Liang, Ding Li 0001, Yao Guo 0001, Xiangqun Chen
LCTES6
2026 Promoguardian: Detecting Promotion Abuse Fraud with Multi-Relation Fused Graph Neural Networks
abstract
As e-commerce platforms develop, fraudulent activities are increasingly emerging, posing significant threats to the security and stability of these platforms. Promotion abuse is one of the fastest-growing types of fraud in recent years and is characterized by users exploiting promotional activities to gain financial benefits from the platform. To investigate this issue, we conduct the first study on promotion abuse fraud in e-commerce platforms MEITUAN. We find that promotion abuse fraud is a group-based fraudulent activity with two types of fraudulent activities: Stocking Up and Cashback Abuse. Unlike traditional fraudulent activities such as fake reviews, promotion abuse fraud typically involves ordinary customers conducting legitimate transactions and these two types of fraudulent activities are often intertwined. To address this issue, we propose leveraging additional information from the spatial and temporal perspectives to detect promotion abuse fraud. In this paper, we introduce PROMOGUARDIAN, a novel multi-relation fused graph neural network that integrates the spatial and temporal information of transaction data into a homogeneous graph to detect promotion abuse fraud. We conduct extensive experiments on real-world data from MEITUAN, and the results demonstrate that our proposed model outperforms state-of-the-art methods in promotion abuse fraud detection, achieving 93.15% precision, detecting 2.1 to 5.0 times more fraudsters, and preventing 1.5 to 8.8 times more financial losses in production environments.
Shaofei Li, Ziqi Zhang 0017, Minyao Hua, Shuli Gao, Zhenkai Liang, Yao Guo 0001, Xiangqun Chen, Ding Li 0001
SP8
2026 PredComp: Predicting Compiler Optimization Options with Multi-stage Learning
abstract
Standard compiler optimization levels, such as -O3 , which provides a fixed optimization strategy for all programs, often fail to deliver the optimal performance. Compiler auto-tuning techniques can deliver substantial speedups, but existing methods present a difficult tradeoff. While dynamic iterative approaches are effective, their requirement for repeated compilation and execution incurs high overhead, which limits their practicality. Conversely, static prediction methods offer a low-overhead alternative. However, they face a vast search space and must comprehensively learn both option-option interactions and option-program feature relationships. To overcome the challenge, we propose PredComp , a novel static framework that leverages the divide and conquer paradigm to predict desired option sets. PredComp decomposes the search space by partitioning options into distinct subspaces based on their relationships, making the prediction problem tractable. It first predicts promising option sub-sets within each subspace, focusing only on intra-subspace option interactions and their preferred program features. Then, it adopts a combination model that aggregates these top-ranked sub-sets, prioritizes inter-subspace option interactions and corresponding features to construct globally desired sets. Experiments on three widely used benchmark suites and one real-world application show that PredComp achieves average speedups of 1.1011× over -O3 with a single prediction. Notably, it achieves performance comparable to dynamic iterative methods while reducing tuning time from hours or days to seconds, thereby making static prediction a practical solution for large-scale and frequently evolving software.
Bingyu Gao, Mengyu Yao, Zhihong Xue, Xiangqun Chen, Ding Li 0001, Yao Guo 0001
ACM Trans. Archit. Code Optim.5
2025 Cost-Efficient Cloud Infrastructure with Hugepage-aware Memory Deduplication
abstract
Optimizing memory cost-efficiency is the top demand for many cloud computing scenarios. Memory deduplication and hugepage are both essential techniques for reducing memory cost and improving efficiency. However, the simultaneous use of memory deduplication and hugepages faces a dillema. Existing approaches either split hugepages into small pages to achieve efficient memory deduplication or ignore redundant portions within hugepages to maintain hugepage performance.
Ruizhe Huang, Xinyu Wang 0043, Zhida An, Hanwen Lei, Peng Jiang 0007, Ziqi Zhang 0017, Ding Li 0001, Yao Guo 0001, Xiangqun Chen, Yuxin Ren 0001, Ning Jia 0004, Xinwei Hu
SoCC9
2025 DPUaudit: DPU-assisted Pull-based Architecture for Near-Zero Cost System Auditing
abstract
System auditing frameworks are crucial for modern data center security, as they record system events to detect intrusions. However, existing software-based auditing frameworks are limited by their high runtime overhead. To address the limitations of software-based frameworks, researchers had proposed a hardware-based auditing framework that offloads log processing to isolated hardware. However, despite using powerful specialized hardware, this approach still suffers from high runtime overhead, which contradicts their efficiency goal. We have identified that the high overhead is due to the pushbased architecture, which involves operating a log sender on the monitored host. Consequently, the existing approach requires heavy software protection mechanisms to secure the log sender, resulting in high runtime overhead.In this paper, we propose a new DPU-assisted pull-based architecture called DPUaudit for hardware-based auditing, which achieves near-zero runtime overhead. Instead of using a log sender, DPUaudit utilizes DPU to actively pull system events from the monitored host. This eliminates the need for heavy mechanisms to handle and safeguard the log sender, achieving highly efficient system auditing. Experimental results show that, on average, DPUaudit only slows down applications on the monitored host by 2.1% for six mainstream data center applications under different workloads, which is at least one order of magnitude smaller than existing approaches, while still ensuring the integrity of audit logs.
Peng Jiang 0007, Hanlin Jiang, Ruizhe Huang, Hanwen Lei, Zhineng Zhong, Yuxin Ren 0001, Ning Jia 0004, Xinwei Hu, Yao Guo 0001, Xiangqun Chen, Ding Li 0001
HPCA11
2025 Grouptuner: Efficient Group-Aware Compiler Auto-tuning
abstract
Modern compilers typically provide hundreds of options to optimize program performance, but users often cannot fully leverage them due to the huge number of options. While standard optimization combinations (e.g., -O3) provide reasonable defaults, they often fail to deliver near-peak performance across diverse programs and architectures. To address this challenge, compiler auto-tuning techniques have emerged to automate the discovery of improved option combinations. Existing techniques typically focus on identifying critical options and prioritizing them during the search to improve efficiency. However, due to limited tuning iterations, the resulting data is often sparse and noisy, making it highly challenging to accurately identify critical options. As a result, these algorithms are prone to being trapped in local optima. To address this limitation, we propose GroupTuner, a group-aware auto-tuning technique that directly applies localized mutation to coherent option groups based on historically best-performing combinations, thus avoiding explicitly identifying critical options. By forgoing the need to know precisely which options are most important, GroupTuner maximizes the use of existing performance data, ensuring more targeted exploration. Extensive experiments demonstrate that GroupTuner can efficiently discover competitive option combinations, achieving an average performance improvement of 12.39% over -O3 while requiring only 77.21% of the time compared to the random search algorithm, significantly outperforming state-of-the-art methods.
Bingyu Gao, Mengyu Yao, Ding Li 0001, Xiangqun Chen, Yao Guo 0001
LCTES6
2025 Predictable and Secure System Auditing for Real-Time Systems
abstract
System auditing frameworks are essential for operating system security as they record system events to support intrusion detection, compliance verification and attack reconstruction. However, existing auditing frameworks fail to meet the stringent requirements of real-time systems, which demand security, predictability, and efficiency. Though current solutions are optimized for security or performance, they do not focus on bounding the worst-case execution time (WCET) and incorporating into response-time analysis (RTA). This paper presents RT-NODROP, a secure and predictable auditing framework tailored for real-time systems. RT-NODROP employs a lightweight threadlet-based architecture to isolate audit events processing, periodically invoking threadlets to simultaneously bound WCET and event residence time. By integrating with real-time schedule, RT-NODROP ensures no event dropping, system efficiency, and predictability. We further develop an overhead-aware RTA and a period selection algorithm to balance security, performance, and schedulability. The evaluations demonstrate that RT-NODROP is superior over state-of-the-art frameworks (Sysdig, OMNILOG, Ellipsis), improving schedulability by$\mathbf{8 0. 1 1 \%, ~} \mathbf{1 1 7. 9 \%}$and$\mathbf{5 1. 0 5 \%}$, respectively. For the latency-intensive application Redis, RT-NODROP achieves up to$\mathbf{7 5. 1 \%}(\mathbf{1 3 8. 8 6 \%}, \mathbf{3 2 4. 6 \%})$higher throughput and$\mathbf{2. 1 9} \times$(3.07x, 5.02x) lower 99.9th percentile tail latency than Sysdig (OMNILOG, Ellipsis) while maintaining a minimum event residence time around 10 ms without event dropping.
Peng Jiang 0007, Fanhang Hu, Ruizhe Huang, Shuomin Xue, Zhaomeng Deng, Yuxin Ren 0001, Ning Jia 0004, Yao Guo 0001, Xiangqun Chen, Ding Li 0001
RTSS9
2025 Query Provenance Analysis: Efficient and Robust Defense Against Query-Based Black-Box Attacks
abstract
Query-based black-box attacks have emerged as a significant threat to machine learning systems, where adversaries can manipulate the input queries to generate adversarial examples that can cause misclassification of the system. To counter these attacks, researchers have proposed Stateful Defense Models (SDMs) such as BlackLight and PIHA, which can reject queries that are “similar” to historical queries. However, recent studies show that existing approaches are vulnerable to a stronger adaptive attack, Oracle-guided Adaptive Rejection Sampling (OARS). OARS can be easily integrated with existing attack algorithms to evade the SDMs by generating queries with fine-tuned direction and step size of perturbations utilizing the leaked decision boundary from the SDMs. In this paper, we propose a novel approach, Query Provenance Analysis (QPA), for defending against query-based black-box attacks robustly (against both non-adaptive and adaptive attacks) and efficiently (in real-time). Our key insight is that, instead of focusing on individual queries, utilizing features from the query sequence (termed query provenance) can distinguish malicious queries from benign queries more effectively. We construct a query provenance graph to capture the relationship between a new query and prior historical queries, and then design efficient algorithms to detect malicious queries based on the query provenance graphs. We evaluate QPA on four datasets against six query-based attacks and compare QPA with state-of-the-art SDM defenses. The results show that QPA outperforms the baselines regarding defense robustness and efficiency on both non-adaptive and adaptive attacks. Specifically, QPA reduces the Attack Success Rate (ASR) of OARS to 4.08%, which is roughly 20× lower than the baselines. Moreover, QPA achieves higher throughput (up to 7.67×) and lower latency (up to 11.09×) than baselines.
Shaofei Li, Haomin Jia, Yao Guo 0001, Xiangqun Chen, Ding Li 0001
SP5
2025 I Can Tell Your Secrets: Inferring Privacy Attributes from Mini-app Interaction History in Super-apps
Yifeng Cai, Mengyu Yao, Xiaoke Zhao, Zhe Liu 0001, Xiangqun Chen, Yao Guo 0001, Ding Li 0001
USENIX Security Symposium9
2025 A survey on EOSIO systems security: vulnerability, attack, and mitigation
Ningyu He, Haoyu Wang 0001, Lei Wu 0012, Xiapu Luo, Yao Guo 0001, Xiangqun Chen
Frontiers Comput. Sci.6
2025 TEESlice: Protecting Sensitive Neural Network Models in Trusted Execution Environments when Attackers Have Pre-Trained Models
abstract
Trusted Execution Environments (TEEs) are used to safeguard on-device models. However, directly employing TEEs to secure the entire DNN model is challenging due to the limited computational speed. Utilizing GPU can accelerate DNN’s computation speed but widely available commercial GPUs usually lack security protection. To this end, scholars introduce TEE-Shielded DNN Partition (TSDP), a method that protects privacy-sensitive weights within TEEs and offloads insensitive weights to GPUs. Nevertheless, current methods do not consider the presence of a knowledgeable adversary who can access abundant publicly available pre-trained models and datasets. This article investigates the security of the existing methods against such a knowledgeable adversary and reveals their inability to fulfill their security promises. Consequently, we introduce a novel partition before training strategy, which effectively separates privacy-sensitive weights from other components of the model. Our evaluation demonstrates that our approach can offer full model protection with a computational cost reduced by a factor of 10. In addition to traditional CNN models, we also demonstrate the scalability to large language models. Our approach can compress the private functionalities of the large language model to lightweight slices and achieve the same level of protection as the shielding-whole-model baseline.
Ding Li 0001, Ziqi Zhang 0017, Mengyu Yao, Yifeng Cai, Yao Guo 0001, Xiangqun Chen
ACM Trans. Softw. Eng. Methodol.6
2025 Not All Exceptions Are Created Equal: Triaging Error Logs in Real-World Enterprises
abstract
Error logs like Java exceptions play a crucial role in diagnosing and resolving errors within the industry. Nonetheless, the extensive logging of Java exceptions may result in exception fatigue in large-scale Java systems at an industrial level, where the frequency of Java exceptions being generated surpasses developers’ ability to manage them effectively. Regrettably, there is a lack of research on the seriousness, prevalence, and solutions to this problem. To close this gap, we first make a comprehensive investigation into the exception fatigue problem within a prominent Internet corporation in China, namely Alibaba, confirming its importance in the industry. Consequently, we introduce a novel solution called ABEL , designed to automatically pinpoint the most relevant exceptions associated with software failures. The key challenge lies in the randomness of exceptions, which prevents existing sequence-based techniques from being effective. To address this challenge, ABEL establishes correlations between Java exceptions and the Key Performance Indicator (KPI) of applications, enabling the identification of exceptions leading to irregularities in KPI. Our evaluation of ABEL across four Java applications and five business KPIs within Alibaba illustrates its capability to pinpoint the primary cause of exception logs with an AC@5 (top-5 accuracy) exceeding 90%, effectively mitigating the exception fatigue problem within Alibaba. Furthermore, it can identify the root-cause exceptions in a real software failure within just 4 minutes, outperforming the manual investigation process by over an hour.
Mengyu Yao, Shaofei Li, Dingyu Yang, Zheshun Wu, Xiaojun Qu, Ziqi Zhang 0017, Ding Li 0001, Yao Guo 0001, Xiangqun Chen
ACM Trans. Softw. Eng. Methodol.10
2024 SeeWasm: An Efficient and Fully-Functional Symbolic Execution Engine for WebAssembly Binaries
abstract
WebAssembly (Wasm), as a compact, fast, and isolation-guaranteed binary format, can be compiled from more than 40 high-level programming languages. However, vulnerabilities in Wasm binaries could lead to sensitive data leakage and even threaten their hosting environments. To identify them, symbolic execution is widely adopted due to its soundness and the ability to automatically generate exploitations. However, existing symbolic executors for Wasm binaries are typically platform-specific, which means that they cannot support all Wasm features. They may also require significant manual interventions to complete the analysis and suffer from efficiency issues as well. In this paper, we propose an efficient and fully-functional symbolic execution engine, named SeeWasm. Compared with existing tools, we demonstrate that SeeWasm supports full-featured Wasm binaries without further manual intervention, while accelerating the analysis by 2 to 6 times. SeeWasm has been adopted by existing works to identify more than 30 0-day vulnerabilities or security issues in well-known C, Go, and SGX applications after compiling them to Wasm binaries.
Ningyu He, Zhehao Zhao, Hanqin Guan, Shuo Peng, Ding Li 0001, Haoyu Wang 0001, Xiangqun Chen, Yao Guo 0001
ISSTA8
2024 Semantic-Enhanced Indirect Call Analysis with Large Language Models
abstract
In contemporary software development, the widespread use of indirect calls to achieve dynamic features poses challenges in constructing precise control flow graphs (CFGs), which further impacts the performance of downstream static analysis tasks. To tackle this issue, various types of indirect call analyzers have been proposed. However, they do not fully leverage the semantic information of the program, limiting their effectiveness in real-world scenarios.
Baijun Cheng, Cen Zhang, Kailong Wang 0001, Ling Shi 0002, Yang Liu 0003, Haoyu Wang 0001, Yao Guo 0001, Ding Li 0001, Xiangqun Chen
ASE9
2024 NODLINK: An Online System for Fine-Grained APT Attack Detection and Investigation
Shaofei Li, Feng Dong 0008, Xusheng Xiao, Haoyu Wang 0001, Fei Shao, Jiedong Chen, Yao Guo 0001, Xiangqun Chen, Ding Li 0001
NDSS8
2024 No Privacy Left Outside: On the (In-)Security of TEE-Shielded DNN Partition for On-Device ML
abstract
On-device ML introduces new security challenges: DNN models become white-box accessible to device users. Based on white-box information, adversaries can conduct effective model stealing (MS) against model weights and membership inference attack (MIA) against training data privacy. Using Trusted Execution Environments (TEEs) to shield on-device DNN models aims to downgrade (easy) white-box attacks to (harder) black-box attacks. However, one major shortcoming of TEEs is the sharply increased latency (up to 50×). To accelerate TEE-shield DNN computation with GPUs, researchers proposed several model partition techniques. These solutions, referred to as TEE-Shielded DNN Partition (TSDP), partition a DNN model into two parts, offloading1the privacy-insensitive part to the GPU while shielding the privacy-sensitive part within the TEE. However, the community lacks an in-depth understanding of the seemingly encouraging privacy guarantees offered by existing TSDP solutions during DNN inference. This paper benchmarks existing TSDP solutions using both MS and MIA across a variety of DNN models, datasets, and metrics. We show important findings that existing TSDP solutions are vulnerable to privacy-stealing attacks and are not as safe as commonly believed. We also unveil the inherent difficulty in deciding the optimal DNN partition configurations, which vary across datasets and models. Based on lessons harvested from the experiments, we present TEESlice, a novel TSDP method that defends against MS and MIA during DNN inference. Unlike existing approaches, TEESlice follows a partition-before-training strategy, which allows for accurate separation between privacy-related weights from public weights. TEESlice delivers the same security protection as shielding the entire DNN model inside TEE (the "upper-bound" security guarantees) with over 10×less overhead (in both experimental and real-world environments) than prior TSDP solutions and no accuracy loss. We make the code and artifacts publicly available on the Internet.
Ziqi Zhang 0017, Yifeng Cai, Yuanyuan Yuan 0001, Ding Li 0001, Yao Guo 0001, Xiangqun Chen
SP8
2024 Detecting Malicious Websites From the Perspective of System Provenance Analysis
abstract
Malicious websites are considered one of the top threats to the modern Internet. Thus, it is critical to effectively detect malicious websites for the security of the Internet. Conventional technologies typically rely on URL blacklists, or static and dynamic code analysis, which are known to have limitations. In order to effectively detect malicious websites, in this paper, we study malicious websites from the perspective of system provenance analysis for the first time. We first conduct a systematic feature engineering study on thousands of benign and malicious websites from the perspective of system provenance data. In our study, we discover eight useful features for malicious website detection. Based on these eight features, we propose ProvWeb, a novel non-intrusive system provenance-based tool, for malicious website detection. In our evaluation, ProvWeb can achieve an F1 score of 93.7% ∼ 99.7% for the four combinations of browsers and OSes (Windows Chrome, Windows Firefox, Linux Chrome, Linux Firefox). This result confirms that the features discovered in provenance graphs are effective in detecting malicious websites.
Peng Jiang 0007, Jifan Xiao, Ding Li 0001, Hongyi Yu, Yao Guo 0001, Xiangqun Chen
IEEE Trans. Dependable Secur. Comput.7
2024 TrajBERT: BERT-Based Trajectory Recovery With Spatial-Temporal Refinement for Implicit Sparse Trajectories
abstract
In the realm of human mobility data analysis, a multitude of constraints result in the publication of sparse, non-uniform implicit trajectories without explicit location information, such as coordinates. Researchers have dedicated substantial efforts towards trajectory recovery, aiming to densify trajectories and gain a more comprehensive understanding of human mobility. However, existing trajectory recovery methods focus on explicit trajectories, and require extensive historical data to capture users' mobility patterns. Nevertheless, implicit trajectories are usually more sparse than explicit trajectories. Addressing these challenges, we propose TrajBERT, an innovative BERT-based trajectory recovery method with spatial-temporal refinement. TrajBERT employs the Transformer encoder to learn mobility patterns bi-directionally and enhances the predictions by cross-stage temporal refinement. Subsequently, we design an output layer with global spatial refinement with a novel spatial-temporal aware loss function. To evaluate the performance of TrajBERT, we conduct a series of experiments on real-world datasets. Remarkably,TrajBERT yields at least 8.2% performance improvement compared to the state-of-the-art trajectory recovery approachs. Furthermore, TrajBERT successfully mitigates the cold start problem commonly experienced with new users lacking historical trajectories. It also shows superior robustness when faced with extremely sparse trajectories, thus demonstrating its potential as a practical tool in the field of human mobility analysis.
Junjun Si, Hanqiu Wang, Li Li 0010, Rongqing Zhang 0001, Bo Tu, Xiangqun Chen
IEEE Trans. Mob. Comput.8
2023 Are we there yet? An Industrial Viewpoint on Provenance-based Endpoint Detection and Response Tools
abstract
Provenance-Based Endpoint Detection and Response (P-EDR) systems are deemed crucial for future Advanced Persistent Threats (APT) defenses. Despite the fact that numerous new techniques to improve P-EDR systems have been proposed in academia, it is still unclear whether the industry will adopt P-EDR systems and what improvements the industry desires for P-EDR systems. To this end, we conduct the first set of systematic studies on the effectiveness and the limitations of P-EDR systems. Our study consists of four components: a one-to-one interview, an online questionnaire study, a survey of the relevant literature, and a systematic measurement study. Our research indicates that all industry experts consider P-EDR systems to be more effective than conventional Endpoint Detection and Response (EDR) systems. However, industry experts are concerned about the operating cost of P-EDR systems. In addition, our research reveals three significant gaps between academia and industry (1) overlooking client-side overhead; (2) imbalancedalarm triage cost and interpretation cost; and (3) excessive server side memory consumption. This paper's findings provide objective data on the effectiveness of P-EDR systems and how much improvements are needed to adopt P-EDR systems in industry.
Feng Dong 0008, Shaofei Li, Peng Jiang 0007, Ding Li 0001, Haoyu Wang 0001, Liangyi Huang, Xusheng Xiao, Jiedong Chen, Xiapu Luo, Yao Guo 0001, Xiangqun Chen
CCS11
2023 Put Your Memory in Order: Efficient Domain-based Memory Isolation for WASM Applications
abstract
Memory corruption vulnerabilities can have more serious consequences in WebAssembly than in native applications. Therefore, we present \tool, the first WebAssembly runtime with memory isolation. Our insight is to use MPK hardware for efficient memory protection in WebAssembly. However, MPK and WebAssembly have different memory models: MPK protects virtual memory pages, while WebAssembly uses linear memory that has no pages. Mapping MPK APIs to WebAssembly causes memory bloating and low running efficiency. To solve this, we propose \acfdilm, which protects linear memory at function-level granularity. We implemented \acdilm into the official WebAssembly runtime to build \tool. Our evaluation shows that \tool can prevent memory corruption in real projects with a 1.77% average overhead and negligible memory cost.
Hanwen Lei, Ziqi Zhang 0017, Peng Jiang 0007, Zhineng Zhong, Ningyu He, Ding Li 0001, Yao Guo 0001, Xiangqun Chen
CCS9
2023 SymGX: Detecting Cross-boundary Pointer Vulnerabilities of SGX Applications via Static Symbolic Execution
abstract
Intel Security Guard Extensions (SGX) have shown effectiveness in critical data protection. Recent symbolic execution-based techniques reveal that SGX applications are susceptible to memory corruption vulnerabilities. While existing approaches focus on conventional memory corruption in ECalls of SGX applications, they overlook an important type of SGX dedicated vulnerability: cross-boundary pointer vulnerabilities. This vulnerability is critical for SGX applications since they heavily utilize pointers to exchange data between secure enclaves and untrusted environments. Unfortunately, none of the existing symbolic execution approaches can effectively detect cross-boundary pointer vulnerabilities due to the lack of an SGX-specific analysis model that properly handles three unique features of SGX applications: Multi-entry Arbitrary-order Execution, Stateful Execution, and Context-aware Pointers. To address such problems, we propose a new analysis model named Global State Transition Graph with Context Aware Pointers (GSTG-CAP) that simulates properties-preserving execution behaviors for SGX applications and drives symbolic execution for vulnerability detection. Based on GSTG-CAP, we build a novel symbolic execution-based vulnerability detector named SYMGX to detect cross-boundary pointer vulnerabilities. According to our evaluation, SYMGX can find 30 0-DAY vulnerabilities in 14 open-source projects, three of which have been confirmed by developers. SYMGX also outperforms two state-of-the-art tools, COIN and TeeRex, in terms of effectiveness, efficiency, and accuracy.
Yuanpeng Wang, Ziqi Zhang 0017, Ningyu He, Zhineng Zhong, Shengjian Guo, Qinkun Bao, Ding Li 0001, Yao Guo 0001, Xiangqun Chen
CCS9
2023 FedSlice: Protecting Federated Learning Models from Malicious Participants with Model Slicing
abstract
Crowdsourcing Federated learning (CFL) is a new crowdsourcing development paradigm for the Deep Neural Network (DNN) models, also called “software 2.0”. In practice, the privacy of CFL can be compromised by many attacks, such as free-rider attacks, adversarial attacks, gradient leakage attacks, and inference attacks. Conventional defensive techniques have low efficiency because they deploy heavy encryption techniques or rely on Trusted Execution Environments (TEEs). To improve the efficiency of protecting CFL from these attacks, this paper proposes FedSlice to prevent malicious participants from getting the whole server-side model while keeping the performance goal of CFL. FedSlice breaks the server-side model into several slices and delivers one slice to each participant. Thus, a malicious participant can only get a subset of the server-side model, preventing them from effectively conducting effective attacks. We evaluate FedSlice against these attacks, and results show that FedSlice provides effective defense: the server-side model leakage is reduced from 100% to 43.45%, the success rate of adversarial attacks is reduced from 100% to 11.66%, the average accuracy of membership inference is reduced from 71.91% to 51.58%, and the data leakage from shared gradients is reduced to the level of random guesses. Besides, FedSlice only introduces less than 2% accuracy loss and about 14% computation overhead. To the best of our knowledge, this is the first paper to discuss defense methods against these attacks to the CFL framework.
Ziqi Zhang 0017, Yuanchun Li 0003, Yifeng Cai, Ding Li 0001, Yao Guo 0001, Xiangqun Chen
ICSE7
2023 A Dual Self-supervised Deep Trajectory Clustering Method
abstract
Trajectory clustering is a cornerstone task in trajectory mining. Sparse and noisy trajectories like Call Detail Records (CDR) have become popular with the rapid development of mobile applications. However, existing trajectory clustering methods' performance is limited on these trajectories. Therefore, we propose a dual Self-supervised Deep Trajectory Clustering (SDTC) method, to optimize trajectory representation and clustering jointly. First, we leverage the BERT model to learn spatial-temporal mobility patterns and incorporate them into the embeddings of location IDs. Second, we fine-tune the BERT model to learn cluster-friendly representations of trajectories by designing a dual self-supervised cluster layer, which improves the intra-cluster similarities and inter-cluster dissimilarities. Third, we conduct extensive experiments with two real-world datasets. Results show that SDTC improves the clustering accuracy by 12.1% (on a noisy and sparse dataset) and 3.8% (on a very sparse dataset) compared with SOTA deep clustering methods.
Junjun Si, Li Li 0010, Bo Tu, Xiangqun Chen, Rongqing Zhang 0001
ISCC6
2023 APIMind: API-driven Assessment of Runtime Description-to-permission Fidelity in Android Apps
abstract
Assessing description-to-permission fidelity is critical for safeguarding personal data accessed through sensitive APIs in Android apps. However, it remains a challenge for existing methods, both static and dynamic. Static methods are either infeasible due to various dynamic features (e.g., code obfuscation, dynamic class loading, and reflection) or too coarse-grained to understand how sensitive APIs collect privacy data under runtime contexts. Existing dynamic methods lack contextual understanding regarding sensitive API calls. For example, they fail to understand which GUI widgets are more likely to trigger sensitive APIs and ignore the preceding UI contexts that could reveal the intention of API calls when analyzing their fidelity.In this paper, we propose an API-driven automated dynamic analysis tool called APIMind for assessing runtime description-to-permission fidelity in Android apps. APIMind can discover sensitive APIs more effectively by utilizing multimodal features to jointly infer the semantics of GUI widgets and leveraging deep networks to automatically learn their relationship based on multifaceted rewards. Then, it could accurately assess description-to-permission fidelity by developing an extended tool that considers dual UI contexts (i.e., preceding and current contexts). We evaluate the accuracy and efficiency of APIMind using 121 real-world apps. Experimental results demonstrate that APIMind can achieve a detection accuracy of 96.1%. Compared to the competitive baseline, APIMind increases efficiency by 43%. In addition, based on our proposed tool, we conduct a large-scale case study of 1013 real Android apps, which reveals the prevalence of several typical inconsistencies and demonstrates the effectiveness of our approach in the wild.
Hanwen Lei, Yuanpeng Wang, Ding Li 0001, Yao Guo 0001, Xiangqun Chen
ISSRE6
2023 ReSPlay: Improving Cross-Platform Record-and-Replay with GUI Sequence Matching
abstract
Record-and-replay is an important testing technique to ensure the quality of mobile applications (apps in short). State-of-the-art record-and-replay approaches are typically based on widget matching, which has shown limited effectiveness, especially on devices with different platforms and resolutions, due to the difficulty in matching widgets with subtle visual differences. Our key observation is that, even if two widgets look similar, the resulting screenshot sequences can still be very different during execution. Thus, instead of matching GUI widgets directly, we are able to find the correct replay actions by comparing the resulting GUI screenshot sequences, which can be better distinguished across different platforms, thus potentially improving the record-and-replay efficiency through GUI exploration and comparison.This paper proposes a general record-and-replay framework called ReSPlay, which leverages a more robust visual feature, GUI sequences, to guide replaying more accurately. ReSPlay pre-trains a deep reinforcement learning model, SDP-Net, offline from random app traces. Specifically, SDP-Net is trained to search a particular path from GUI transition graphs to learn an optimal policy to locate the target operation positions by maximizing the possibilities to reach the target GUI sequence. Finally, the trained SDP-Net is used to search for potential event traces with high rewards and replicate them on the target device for replay. We evaluate our proposed framework on multiple real devices. Experimental results show that the overall average replay accuracy of ReSPlay on devices across different OSes, GUI styles, and resolutions is 28.12% higher than the state-of-the-art baselines.
Linna Wu, Yuanchun Li 0003, Ziqi Zhang 0017, Hanwen Lei, Ding Li 0001, Yao Guo 0001, Xiangqun Chen
ISSRE8
2023 Eunomia: Enabling User-Specified Fine-Grained Search in Symbolically Executing WebAssembly Binaries
abstract
Although existing techniques have proposed automated approaches to alleviate the path explosion problem of symbolic execution, users still need to optimize symbolic execution by applying various searching strategies carefully. As existing approaches mainly support only coarse-grained global searching strategies, they cannot efficiently traverse through complex code structures. In this paper, we propose Eunomia, a symbolic execution technique that supports fine-grained search with local domain knowledge. Eunomia uses Aes, a DSL that lets users specify local searching strategies for different parts of the program. Eunomia also isolates the context of variables for different local searching strategies, avoiding conflicts. We implement Eunomia for WebAssembly, which can analyze applications written in various languages. Eunomia is the first symbolic execution engine that supports the full features of WebAssembly. We evaluate Eunomia with a microbenchmark suite and six real-world applications. Our evaluation shows that Eunomia improves bug detection by up to three orders of magnitude. We also conduct a user study that shows the benefits of using Aes. Moreover, Eunomia verifies six known bugs and detects two new zero-day bugs in Collections-C.
Ningyu He, Zhehao Zhao, Yubin Hu 0003, Shengjian Guo, Haoyu Wang 0001, Guangtai Liang, Ding Li 0001, Xiangqun Chen, Yao Guo 0001
ISSTA9
2023 How Android Apps Break the Data Minimization Principle: An Empirical Study
abstract
The Data Minimization Principle is crucial for protecting individual privacy. However, existing Android runtime permissions do not guarantee this principle. Moreover, the lack of an automatic enforcement mechanism leads to uncertainty as to whether apps strictly comply with this principle. To bridge this gap, we conduct the first systematic empirical study on violations of the Data Minimization Principle and design a new enforcement tool called GUIMind to detect them. GUIMind first utilizes a reinforcement learning model to explore app activities and monitor access to sensitive APIs that require sensitive permissions, and then it leverages an existing tool to detect such violations. We evaluate the performance of GUIMind using 120 real-world Android apps. The results indicate that GUIMind can achieve a detection accuracy of 96.1%, effectively accelerating the empirical study. Our empirical research is mainly focused on the prevalence of violations, the responses of administrators to violations, and the potential factors and characteristics that lead to violations, such as typical violations, app categories, and personal data types. Our study reveals that 83.5% of apps contain at least one privacy violation, with health apps being the most severe. In addition, telephony information is the most commonly leaked personal data type, accounting for 71.1%. Finally, we randomly selected 60 non-compliant apps for reporting to the administrator, whose responses confirm the effectiveness of our approach.
Hanwen Lei, Yuanpeng Wang, Ding Li 0001, Yao Guo 0001, Xiangqun Chen
ASE6
2023 Auditing Frameworks Need Resource Isolation: A Systematic Study on the Super Producer Threat to System Auditing and Its Mitigation
Peng Jiang 0007, Ruizhe Huang, Ding Li 0001, Yao Guo 0001, Xiangqun Chen, Jianhai Luan, Yuxin Ren 0001, Xinwei Hu
USENIX Security Symposium5
2023 Beyond Fine-Tuning: Efficient and Effective Fed-Tuning for Mobile/Web Users
abstract
Fine-tuning is a typical mechanism to achieve model adaptation for mobile/web users, where a model trained by the cloud is further retrained to fit the target user task. While traditional fine-tuning has been proved effective, it only utilizes local data to achieve adaptation, failing to take advantage of the valuable knowledge from other mobile/web users. In this paper, we attempt to extend the local-user fine-tuning to multi-user fed-tuning with the help of Federated Learning (FL). Following the new paradigm, we propose EEFT, a framework aiming to achieve Efficient and Effective Fed-Tuning for mobile/web users. The key idea is to introduce lightweight but effective adaptation modules to the pre-trained model, such that we can freeze the pre-trained model and just focus on optimizing the modules to achieve cost reduction and selective task cooperation. Extensive experiments on our constructed benchmark demonstrate the effectiveness and efficiency of the proposed framework.
Yifeng Cai, Hongzhe Bi, Ziqi Zhang 0015, Ding Li 0001, Yao Guo 0001, Xiangqun Chen
WWW7
2022 ReMoS: Reducing Defect Inheritance in Transfer Learning via Relevant Model Slicing
abstract
Transfer learning is a popular software reuse technique in the deep learning community that enables developers to build custom models (students) based on sophisticated pretrained models (teachers). However, like vulnerability inheritance in traditional software reuse, some defects in the teacher model may also be inherited by students, such as well-known adversarial vulnerabilities and backdoors. Reducing such defects is challenging since the student is unaware of how the teacher is trained and/or attacked. In this paper, we propose ReMoS, a relevant model slicing technique to reduce defect inheritance during transfer learning while retaining useful knowledge from the teacher model. Specifically, ReMoS computes a model slice (a subset of model weights) that is relevant to the student task based on the neuron coverage information obtained by profiling the teacher model on the student task. Only the relevant slice is used to finetune the student model, while the irrelevant weights are retrained from scratch to minimize the risk of inheriting defects. Our experiments on seven DNN defects, four DNN models, and eight datasets demonstrate that ReMoS can reduce inherited defects effectively (by 63% to 86% for CV tasks and by 40% to 61% for NLP tasks) and efficiently with minimal sacrifice of accuracy (3% on average).
Ziqi Zhang 0017, Yuanchun Li 0003, Jindong Wang 0001, Ding Li 0001, Yao Guo 0001, Xiangqun Chen, Yunxin Liu 0001
ICSE7
2022 City-Scale Fingerprint Positioning Framework based on MDT Data
abstract
Mobile positioning plays an essential role in smart city services. This paper proposes a fingerprint positioning framework based on massive Minimization of Drive Test (MDT) data to provide accurate and efficient city-scale positioning without additional equipment and measures. First, a multi-level fingerprint construction method is proposed using the Timing Advance (TA), Reference Signal Receiving Power (RSRP), and Reference Signal Receiving Quality (RSRQ) of the serving cell and neighboring cell. Then, an adaptive online fingerprint matching method is employed to extract and match online data fingerprints. Experiments show that the median positioning error is 29.97 meters with city-scale MDT data. It outperforms the reported accuracy of the state-of-the-art fingerprint positioning method.
Junjun Si, Zejiang Chen, Bo Tu, Shuaifu Dai, Xiangqun Chen
ISCC5
2022 Toward a Systematic Survey on Wearable Computing for Education Applications
abstract
Technology is gradually being incorporated as an integral part of education. This technology includes the now pervasive presence of the Internet, but also the inclusion of devices that are worn by teachers and learners. Wearable devices have greatly facilitated data acquisition of educational subjects. This article provides an overview of what and how wearable devices have been used in education and in what contexts. Summaries of existing research in this area are organized according to a three-layered data framework. We begin by presenting the technical characteristics of four types of wearable devices, and comparing the data from them with physiological data and behavioral data. Based on this, the preprocessing and analysis methods of these two types of data are discussed. Following that, we identified the six main applications that wearables and data analytics are enabling education. Furthermore, three key challenges of using wearable devices have been identified. We suggest that future research should be enriched from equipment innovation, application scenarios, and feedback methods in order to make wearable devices more widely used in a production environment, and augment the analysis performed to improve teaching, learning, or the educational context where it occurs.
Wei Gao 0037, Xiangqun Chen, Qing Li 0045
IEEE Internet Things J.4
2022 A Graph-Based Temporal Attention Framework for Multi-Sensor Traffic Flow Forecasting
abstract
Accurate spatio-temporal traffic forecasting serves as the basis of dynamic strategy and applications for intelligent transportation systems, which is of great practical significance for improving traffic safety and mitigating road congestion. Recently, deep learning methods such as convolutional neural networks (CNN) have been applied to traffic flow forecasting, which exhibits better performance than conventional methods. However, these CNN-based methods typically learn traffic as images to model spatial correlation, which is only applicable to Euclidean grid map data rather than non-Euclidean multi-sensor data. To address this problem, we propose a graph-based temporal attention framework GTA, which considers both spatial and temporal correlation, to forecast traffic flow based on data collected from multiple sensors. More specifically, GTA can better capture spatial dependencies leveraging graph embedding techniques on sensor networks because it preserves more details in the algorithms. We also introduce an attention mechanism to adaptively identify the relations among temporal submodules. Spatio-temporal dependencies are more effectively and comprehensively integrated due to the full use of the topological properties of transportation networks. We evaluate GTA with a large-scale traffic dataset from England and enhance it with topology information. The experimental results show that our approach outperforms several state-of-the-art baselines.
Yao Guo 0001, Peize Zhao, Chuanpan Zheng, Xiangqun Chen
IEEE Trans. Intell. Transp. Syst.5
2021 TransTailor: Pruning the Pre-trained Model for Improved Transfer Learning
abstract
The increasing of pre-trained models has significantly facilitated the performance on limited data tasks with transfer learning. However, progress on transfer learning mainly focuses on optimizing the weights of pre-trained models, which ignores the structure mismatch between the model and the target task. This paper aims to improve the transfer performance from another angle - in addition to tuning the weights, we tune the structure of pre-trained models, in order to better match the target task. To this end, we propose TransTailor, targeting at pruning the pre-trained model for improved transfer learning. Different from traditional pruning pipelines, we prune and fine-tune the pre-trained model according to the target-aware weight importance, generating an optimal sub-model tailored for a specific target task. In this way, we transfer a more suitable sub-structure that can be applied during fine-tuning to benefit the final performance. Extensive experiments on multiple pre-trained models and datasets demonstrate that TransTailor outperforms the traditional pruning methods and achieves competitive or even better performance than other state-of-the-art transfer learning methods while using a smaller model. Notably, on the Stanford Dogs dataset, TransTailor can achieve 2.7% accuracy improvement over other transfer methods with 20% fewer FLOPs.
Yifeng Cai, Yao Guo 0001, Xiangqun Chen
AAAI4
2021 Dependency-aware Form Understanding
abstract
Form understanding is an important task in many fields such as software testing, AI assistants, and improving accessibility. One key goal of understanding a complex set of forms is to identify the dependencies between form elements. However, it remains a challenge to capture the dependencies accurately due to the diversity of UI design patterns and the variety in development experiences. In this paper, we propose a deep-learning-based approach called DependEX, which integrates convolutional neural networks (CNNs) and transformers to help understand dependencies within forms. DependEX extracts semantic features from UI images using CNN-based models, captures contextual patterns using a multilayer transformer encoder module, and models dependencies between form elements using two embedding layers. We evaluate DependEX with a large-scale dataset from mobile Web applications. Experimental results show that our proposed model achieves over 92% accuracy in identifying dependencies between UI elements, which significantly outperforms other competitive methods, especially for heuristic-based methods. We also conduct case studies on automatic form filling and test case generation from natural language (NL) instructions, which demonstrates the applicability of our approach.
Yuanchun Li 0003, Weixiang Yan, Yao Guo 0001, Xiangqun Chen
ISSRE5
2021 PFA: Privacy-preserving Federated Adaptation for Effective Model Personalization
abstract
Federated learning (FL) has become a prevalent distributed machine learning paradigm with improved privacy. After learning, the resulting federated model should be further personalized to each different client. While several methods have been proposed to achieve personalization, they are typically limited to a single local device, which may incur bias or overfitting since data in a single device is extremely limited. In this paper, we attempt to realize personalization beyond a single client. The motivation is that during the FL process, there may exist many clients with similar data distribution, and thus the personalization performance could be significantly boosted if these similar clients can cooperate with each other. Inspired by this, this paper introduces a new concept called federated adaptation, targeting at adapting the trained model in a federated manner to achieve better personalization results. However, the key challenge for federated adaptation is that we could not outsource any raw data from the client during adaptation, due to privacy concerns. In this paper, we propose PFA, a framework to accomplish Privacy-preserving Federated Adaptation. PFA leverages the sparsity property of neural networks to generate privacy-preserving representations and uses them to efficiently identify clients with similar data distributions. Based on the grouping results, PFA conducts an FL process in a group-wise way on the federated model to accomplish the adaptation. For evaluation, we manually construct several practical FL datasets based on public datasets in order to simulate both the class-imbalance and background-difference conditions. Extensive experiments on these datasets and popular model architectures demonstrate the effectiveness of PFA, outperforming other state-of-the-art methods by a large margin while ensuring user privacy. We will release our code at: https://github.com/lebyni/PFA.
Yao Guo 0001, Xiangqun Chen
WWW3
2020 PrTaurus: An Availability-Enhanced EMR Service on Preemptible Cloud Instances
abstract
EMR (Elastic Map Reduce) is a service provided by mainstream cloud vendors for data processing users to directly obtain well-managed Hadoop YARN clusters on the cloud. Preemptible instance is a kind of cloud server that is cheap but is likely to be reclaimed by cloud vendors suddenly. Running EMR clusters on preemptible instances relies on YARN's own fault-tolerance, which is limited. In this paper, we present PrTaurus as an availability-enhanced EMR service on preemptible instances. PrTaurus integrates a system-level checkpoint capability based on Docker into YARN to further improve its fault-tolerance. In addition, PrTaurus's scheduling strategy takes advantage of Alibaba Cloud's one-hour protection policy. Furthermore, a new method that comprehensively considers cost-efficiency, preemption risk and overhead is proposed to select cluster instances. We evaluated PrTaurus through simulations on real-world workload and instance price traces. Experimental results show that compared with the existing EMR clusters running on preemptible instances, PrTaurus significantly reduces cost (13.0%-74.6%), instance preemptions (60.3%-88.9%), and task preemptions (86.0 % - 98.6 %).
Junming Ma, Yan Li 0067, Xiangqun Chen, Donggang Cao
ICWS3
2020 Dynamic slicing for deep neural networks
abstract
Program slicing has been widely applied in a variety of software engineering tasks. However, existing program slicing techniques only deal with traditional programs that are constructed with instructions and variables, rather than neural networks that are composed of neurons and synapses. In this paper, we introduce NNSlicer, the first approach for slicing deep neural networks based on data-flow analysis. Our method understands the reaction of each neuron to an input based on the difference between its behavior activated by the input and the average behavior over the whole dataset. Then we quantify the neuron contributions to the slicing criterion by recursively backtracking from the output neurons, and calculate the slice as the neurons and the synapses with larger contributions. We demonstrate the usefulness and effectiveness of NNSlicer with three applications, including adversarial input detection, model pruning, and selective model protection. In all applications, NNSlicer significantly outperforms other baselines that do not rely on data flow analysis.
Ziqi Zhang 0017, Yuanchun Li 0003, Yao Guo 0001, Xiangqun Chen, Yunxin Liu 0001
ESEC/SIGSOFT FSE4
2019 How to Find It Better?: Cross-Learning for WeChat Mini Programs
abstract
WeChat Mini Program is a lightweight app relying on the WeChat client, which can be accessed directly from the search list without downloading and installing. Retrieval and ranking for the Mini Programs differ from traditional web search in two sides. On the one hand, as the search queries are often short and most Mini Programs contain few useful textual information, it is hard to retrieve when the user input is inaccurate. On the other hand, without user scoring and rating system like App Store and Google Play, it is hard to rank relatively better results in a more advanced position. In this paper, we propose a Cross-Learning strategy to improve the search experience, where the semantics of queries and Mini Programs are represented not by itself, but by each other. We treat the search task as an extreme multi-label classification problem where the queries are inputs and the Mini Programs are labels. We propose a N-Gram self-attention query encoder to capture the search intention behind these short queries, and carefully design the label selection strategy based on user behavior to rank higher quality Mini Programs in higher positions. Our model outperforms some state-of-the-art baselines in the offline environment, and brought improvement to our actual business in the online A/B Test, which proves the practical significance of our work.
Xiangqun Chen
CIKM5
2019 DCStore: A Deduplication-Based Cloud-of-Clouds Storage Service
abstract
The increasing popularity of cloud storage is leading many organizations to move their data into the cloud. However, putting all data in one cloud causes problems such as vendor lock-in, increased service costs, and data availability. In this paper, we introduce DCStore, a Cloud-of-Clouds storage service designed for an organization to outsource their data into the clouds. To achieve the goal of cost-efficient and high-available, we combine three key techniques. First, DCStore eliminates the redundant data at client-side to save storage cost via application-aware chunking method. Second, DCStore uses an inner-chunk based erasure coding scheme to distribute unique chunks across multiple clouds for high availability. Finally, a container-based share management strategy is used for performance optimization. Our experimental evaluations show that DCStore can improve the performance and cost efficiency significantly, compared with existing Cloud-of-Clouds storage systems.
Bo An 0003, Yan Li 0067, Junming Ma, Gang Huang 0001, Xiangqun Chen, Donggang Cao
ICWS5
2019 Humanoid: A Deep Learning-Based Approach to Automated Black-box Android App Testing
abstract
Automated input generators must constantly choose which UI element to interact with and how to interact with it, in order to achieve high coverage with a limited time budget. Currently, most black-box input generators adopt pseudo-random or brute-force searching strategies, which may take very long to find the correct combination of inputs that can drive the app into new and important states. We propose Humanoid, an automated black-box Android app testing tool based on deep learning. The key technique behind Humanoid is a deep neural network model that can learn how human users choose actions based on an app's GUI from human interaction traces. The learned model can then be used to guide test input generation to achieve higher coverage. Experiments on both open-source apps and market apps demonstrate that Humanoid is able to reach higher coverage, and faster as well, than the state-of-the-art test input generators. Humanoid is open-sourced at https://github.com/yzygitzh/Humanoid and a demo video can be found at https://youtu.be/PDRxDrkyORs.
Yuanchun Li 0003, Yao Guo 0001, Xiangqun Chen
ASE4
2019 WealthAdapt: A General Network Adaptation Framework for Small Data Tasks
abstract
In this paper, we propose a general network adaptation framework, namely WealthAdapt, to effectively adapt a large network for small data tasks, with the assistance of a wealth of related data. While many existing algorithms have proposed network adaptation techniques for resource-constrained systems, they typically implement network adaptation based on a large dataset and do not perform well when facing small data tasks. Because small data have poor feature expression ability, it may result in incorrect filter selection and overfitting during fine-tuning in the network adaptation process. In WealthAdapt, we first expand the target small data task with the wealth of big data, before we perform network adaptation, in order to enrich the features and improve the fine-tuning performance during adaptation. We formally establish network adaptation for small data tasks as an optimization problem and solve it through two main techniques:model-based fast selection andwealth-incorporated iteration adaptation. Experimental results demonstrate that our framework is applicable to both the vanilla convolutional network VGG-16 and more complex modern architecture ResNet-50, outperforming several state-of-the-art network adaptation pipelines on multiple visual classification tasks includinggeneral object recognition, fine-grained object recognition andscene recognition.
Yao Guo 0001, Xiangqun Chen
ACM Multimedia3
2018 Traffic Danger Recognition With Surveillance Cameras Without Training Data
abstract
We propose a traffic danger recognition model that works with arbitrary traffic surveillance cameras to identify and predict car crashes. There are too many cameras to monitor manually. Therefore, we developed a model to predict and identify car crashes from surveillance cameras based on a 3D reconstruction of the road plane and prediction of trajectories. For normal traffic, it supports real-time proactive safety checks of speeds and distances between vehicles to provide insights about possible high-risk areas. We achieve good prediction and recognition of car crashes without using any labeled training data of crashes. Experiments on the BrnoCompSpeed dataset show that our model can accurately monitor the road, with mean errors of 1.80% for distance measurement, 2.77 km/h for speed measurement, 0.24 m for car position prediction, and 2.53 km/h for speed prediction.
Lijun Yu, Xiangqun Chen, Alex Hauptmann 0001
AVSS3
2018 Automated Extraction of Personal Knowledge from Smartphone Push Notifications
abstract
Personalized services are in need of a rich and powerful personal knowledge base, i.e. a knowledge base containing information about the user. This paper proposes an approach to extracting personal knowledge from smartphone push notifications, which are used by mobile systems and apps to inform users of a rich range of information. Our solution is based on the insight that most notifications are formatted using templates, while knowledge entities can be usually found within the parameters to the templates. As defining all the notification templates and their semantic rules are impractical due to the huge number of notification templates used by potentially millions of apps, we propose an automated approach for personal knowledge extraction from push notifications. We first discover notification templates through pattern mining, then use machine learning to understand the template semantics. Based on the templates and their semantics, we are able to translate notification text into knowledge facts automatically. Users' privacy is preserved as we only need to upload the templates to the server for model training, which do not contain any personal information. According to experiments with about 120 million push notifications from 100,000 smartphone users, our system is able to extract personal knowledge accurately and efficiently.
Yuanchun Li 0003, Yao Guo 0001, Xiangqun Chen, Yuvraj Agarwal, Jason I. Hong
IEEE BigData4
2018 Towards Light-Weight Deep Learning Based Malware Detection
abstract
The explosive amount of malware continues threating the security of operating systems and networks. Traditional malware detection approaches fail to meet the requirements of detecting polymorphic and new samples. Existing neural network based detection approaches performs better, but consuming much more time in both feature extraction and training. In this paper, we propose a light-weight PC malware detection system which is based on deep convolutional neural network (CNN). The raw inputs of our system are sequences of grouped instructions, which were generated by our Instruction Analyzer in according to different functionalities of the instructions. The network will automatically learn features of malware from the grouped instruction sequences. The experiment results suggest that in a large dataset which contains roughly 70,000 samples, our detection system can achieve an overall accuracy of 95\%. The training time of our system with single convolutional layer was only about 10 hours, which is one order of magnitude less than traditional methods.
Zeliang Kan, Haoyu Wang 0001, Guoai Xu, Yao Guo 0001, Xiangqun Chen
COMPSAC (1)5
2018 DistGear: A Lightweight Event-Driven Framework for Developing Distributed Applications
abstract
The emergence of cloud computing makes it easier for enterprises and individual users to obtain distributed computing resources. At the same time, developing distributed systems is still complex and difficult. In order to reduce the complexity in distributed application development, a series of techniques such as RPC, message-oriented middleware have been proposed and widely used. These techniques hide the underlying communicating details but require developers to take control of coordination among different machines for complex distributed tasks. This paper proposes a lightweight event-driven framework called DistGear to simplify the development of complex task processing in distributed applications. DistGear is a layer of software built on top of communication facilities for providing developers with a programming abstraction to easily handle distributed tasks and control machine coordination in tasks. We implement DistGear based on coroutines and message communication. Experience with DistGear in a simplified real-world system suggests it is effective. Performance evaluation also demonstrates DistGear's good concurrency and low resources consumption.
Junming Ma, Bo An 0003, Xiangqun Chen, Donggang Cao
COMPSAC (1)3
2018 Shaping program repair space with existing patches and similar code
abstract
Automated program repair (APR) has great potential to reduce bug-fixing effort and many approaches have been proposed in recent years. APRs are often treated as a search problem where the search space consists of all the possible patches and the goal is to identify the correct patch in the space. Many techniques take a data-driven approach and analyze data sources such as existing patches and similar source code to help identify the correct patch. However, while existing patches and similar code provide complementary information, existing techniques analyze only a single source and cannot be easily extended to analyze both.
Jiajun Jiang, Yingfei Xiong 0001, Hongyu Zhang 0002, Xiangqun Chen
ISSTA5
2018 What's inside my app?: understanding feature redundancy in mobile apps
abstract
As the number of mobile apps increases rapidly, many users may install dozens of, or even hundreds of, apps on a single smartphone. However, many apps on the same phone may contain similar or even the same feature, resulting in feature redundancy. For example, multiple apps may check weather forecast for the user periodically. Feature redundancy may cause many undesirable side-effects such as consuming extra CPU resources and network traffic. This paper proposes a method to identify common features within an app, and evaluated it on over four thousand popular apps. Experiments on a list of apps installed on actual smartphones show that the extent of feature redundancy is very high. We found that more than 85% of user smartphones contain redundant features, while in extreme cases, some smartphones may contain dozens of apps with the same feature. In addition, our user surveys found out that about half of the redundant features are undesirable from the end users' perspective, which indicates that feature redundancy has become an important issue that needs to be investigated further.
Yao Guo 0001, Yuanchun Li 0003, Xiangqun Chen
ICPC4
2018 Inferring UI States of Mobile Applications Through Power Side Channel Exploitation
Yao Guo 0001, Junming Ma, Xiangqun Chen
SecureComm (1)4
2018 Building application-specific operating systems: a profile-guided approach
Pengfei Yuan, Yao Guo 0001, Lu Zhang 0023, Xiangqun Chen, Hong Mei 0001
Sci. China Inf. Sci.4
2017 AgileRabbit: A Feedback-Driven Offloading Middleware for Smartwatch Apps
abstract
With the rapid development of wearable devices such as smartwatches, we are brought to a new era of wearable computing. Due to limited computational capability, storage, and battery capacity, wearable devices can hardly execute computation-intensive tasks. The mainstream approach to overcoming these limitations is computation offloading, i.e., offloading the tasks to mobile devices or the remote cloud servers. However, computation offloading cannot improve performance or save power consumption under all conditions. For example, offloading may not be worth in the case of very poor network conditions. To address the issue, in this paper, we propose AgileRabbit, a feedback-driven middleware of computation offloading for smartwatch apps. We design an offloading decision algorithm using the feedback data with a given objective i.e., minimizing the task completion time, or minimizing the total power consumption of smartwatches and mobile devices. With the assistance of AgileRabbit, computation-intensive tasks in smartwatch apps can be well scheduled and assigned to the proper computation node. We implement a speech recognition application on Android Wear platform and deploy it on AgileRabbit to validate the effectiveness of our approach. Evaluation results show that AgileRabbit can significantly improve the performance and save power consumption while incurring small overheads.
Meihua Yu, Yun Ma 0002, Xuanzhe Liu, Gang Huang 0001, Xiangqun Chen
Internetware5
2017 E-Spector: Online energy inspection for Android applications
abstract
Energy consumption is one of the most important aspects of mobile apps. During energy testing, it is important for developers to understand not only the energy consumption rate of an app, but also why energy is consumed. However, existing energy testing tools are more concerned about the accuracy of energy estimation, while typically not providing explanations on why and how exactly energy has been consumed. This paper presents E-Spector, an online energy inspection method for Android apps, which can not only visualize the energy consumption of an app in an instant online manner, but also can tell what happened behind each energy hotspot on the energy curve. E-Spector relies on static analysis and app instrumentation to collect the activities from an app execution in real-time. Then it presents the activities on an instant energy curve, such that the user can easily tell what happened behind each energy spike. Experimental result shows that the energy estimation error of E-Spector is less than 10% and its overhead on energy consumption is about 4%. We also show case studies to demonstrate the applicability and effectiveness of E-Spector in energy monitoring, analysis and bug inspection.
Chengke Wang, Yao Guo 0001, Xiangqun Chen
ISLPED4
2017 Looxy: Web Access Optimization for Mobile Applications with a Local Proxy
abstract
Efficient web caching and prefetching can help optimize mobile application web accesses through eliminating network traffic and reducing human perceived latency. However, due to limited storage and computation resources, it is not practical to implement web caching and prefetching mechanisms on mobile devices such as smartphones. This paper proposes Looxy, a mechanism to optimize mobile application web accesses by offloading the caching and prefetching functionalities to a local proxy. As a result, mobile applications can benefit from these optimizations without incurring extra computation and storage overhead on smartphones. Looxy does not require modifications to the applications or mobile operating systems, making it applicable to different mobile operating systems and devices. Experimental results show that Looxy can save about 20% of Internet traffic with real user behaviors and improve application response speed by about 1.8X.
Yao Guo 0001, Mengxin Liu, Xiangqun Chen
VTC Spring3
2017 FreeNavi: Landmark-Based Mapless Indoor Navigation Based on WiFi Fingerprints
abstract
Although a number of indoor navigation approaches have been proposed, most either require prior knowledge on floor plans, or relying on extra sensors or images, to provide accurate indoor localization and navigation. This paper presents FreeNavi, a landmark-based indoor navigation algorithm that leverages only WiFi signals to direct users in sophisticated indoor environments without prior device deployment or floor plans. FreeNavi takes advantage of human intelligence as an important input to locate and navigate users based on landmarks. With WiFi fingerprints collected at landmarks and walking traces collected in a crowdsourced manner, FreeNavi is able to create a virtual map connecting landmarks with each other. During navigation, FreeNavi produces human understandable directions based on landmarks in the virtual map. Evaluation result shows that FreeNavi can build mostly correct maps and provide efficient directions to users despite relying only on WiFi signals.
Yao Guo 0001, Xiangqun Chen
VTC Spring3
2017 An Explorative Study of the Mobile App Ecosystem from App Developers' Perspective
abstract
With the prevalence of smartphones, app markets such as Apple App Store and Google Play has become the center stage in the mobile app ecosystem, with millions of apps developed by tens of thousands of app developers in each major market. This paper presents a study of the mobile app ecosystem from the perspective of app developers. Based on over one million Android apps and 320,000 developers from Google Play, we analyzed the Android app ecosystem from different aspects. Our analysis shows that while over half of the developers have released only one app in the market, many of them have released hundreds of apps. We classified developers into different groups based on the number of apps they have released, and compared their characteristics. Specially, we have analyzed the group of aggressive developers who have released more than 50 apps, trying to understand how and why they create so many apps. We also investigated the privacy behaviors of app developers, showing that some developers have a habit of producing apps with low privacy ratings. Our study shows that understanding the behavior of mobile developers can be helpful to not only other app developers, but also to app markets and mobile users.
Haoyu Wang 0001, Zhe Liu 0001, Yao Guo 0001, Xiangqun Chen, Miao Zhang 0011, Guoai Xu, Jason I. Hong
WWW4
2016 Self-Adaptive Step Counting on Smartphones under Unrestricted Stepping Modes
abstract
Pedometer apps on smartphones and wearable electronics are increasingly popular nowadays, as they are widely used for health monitoring and location-based systems. Most pedometer apps are based on inertial sensors and each step counting algorithm works precisely under restricted stepping modes because steps are detected and validated after comparing their parameters to pre-determined optimal values related to particular conditions. In this paper, we propose self-adaptive step counting in order to improve step counting accuracy under unrestricted stepping modes on smartphones. Based on our human stepping model, we propose self-adaptive method that can detect new steps by monitoring vertical acceleration and validate new steps by comparing to self-adaptive values, which are adjusted dynamically after each step occurs. We show the flexibility of the proposed approach by incorporating it into two existing step counting algorithms WPD and DTW. We also propose a new stepping cycle recognition (SCR) algorithm that is self-adaptive and performs the best under variant stepping modes. With experiments under different stepping modes including fixed and variant modes, we show that self-adaptive methods perform significantly better compared to original methods using fixed optimal values.
Yao Guo 0001, Xiangqun Chen
COMPSAC3
2016 PERUIM: understanding mobile application privacy with permission-UI mapping
abstract
Current mobile operating systems such as Android employ the permission-based access control mechanism, but it is difficult for users to understand how and why the permissions are used within a particular application. This paper introduces permission-UI mapping as an easy-to-understand representation to illustrate how permissions are used by different UI components within a given application. Connecting UI components to permissions helps users to understand the purpose of permission requests and also makes it possible to illustrate permission requests in a fine-grained manner. We propose PERUIM to extract the permission-UI mapping from an application based on both dynamic and static analysis, and represent the analysis results with a graphical representation. Experiments on popular mobile applications demonstrate the accuracy and applicability of the proposed approach.
Yuanchun Li 0003, Yao Guo 0001, Xiangqun Chen
UbiComp3
2015 Reevaluating Android Permission Gaps with Static and Dynamic Analysis
abstract
Recent studies on the Android permission system have found that there exists a permission gap between the requested permissions and permissions actually used in an Android app. However, current approaches face some challenges when detecting such permission gaps in Android apps due to the limitation of static analysis techniques. This paper proposes a novel approach to detect permission gaps in Android apps and determine the precise set of permissions that an app needs to run correctly. Our approach includes a static analysis technique to extract permission usage information from API invocations, and a dynamic testing technique to test and monitor the runtime permission usage behaviors of apps. By combining static analysis and dynamic testing, our approach can detect significantly more permission usage information compared to static analysis, indicating that our approach could improve the detection accuracy and reduce the false positives in permission gap detection. We have implemented a prototype to study more than 1,000 popular apps from Google Play. The results show that our approach could detect on average 30% more permissions that are used in apps, while more than 8% of the overprivileged apps detected by previous approaches are false positives.
Haoyu Wang 0001, Yao Guo 0001, Guangdong Bai, Xiangqun Chen
GLOBECOM5
2015 A Study on Power Side Channels on Mobile Devices
abstract
Power side channel is a very important category of side channels, which can be exploited to steal confidential information from a computing system by analyzing its power consumption. In this paper, we demonstrate the existence of various power side channels on popular mobile devices such as smartphones. Based on unprivileged power consumption traces, we present a list of real-world attacks that can be initiated to identify running apps, infer sensitive UIs, guess password lengths, and estimate geo-locations. These attack examples demonstrate that power consumption traces can be used as a practical side channel to gain various confidential information of mobile apps running on smartphones. Based on these power side channels, we discuss possible exploitations and present a general approach to exploit a power side channel on an Android smartphone, which demonstrates that power side channels pose imminent threats to the security and privacy of mobile users. We also discuss possible countermeasures to mitigate the threats of power side channels.
Yao Guo 0001, Xiangqun Chen, Hong Mei 0001
Internetware3
2015 Fixing sensor-related energy bugs through automated sensing policy instrumentation
abstract
As mobile applications (apps) become more and more complex, many apps contain various energy bugs, which may cause energy wastes that might reduce the battery life to as short as several hours. Among them, sensor-related bugs such as sensor data underutilization is one of the most common energy bugs. Instead of trying to detect these energy bugs, this paper proposes a method to fix sensor data underutilization automatically through instrumentation of existing apps. App-specific energy-aware sensing policies can be written to the apps via an automated instrumentation process, which can also be customized by users if needed. The proposed technique is easy to apply as it does not need to modify the operating system or the apps. At the same time, it also works for existing legacy apps, which makes it practical and feasible for a wide-range of mobile apps. Experimental results on popular Android apps show that we are able to achieve significant energy savings through automated instrumentation and rebuilding the targeted apps.
Yuanchun Li 0003, Yao Guo 0001, Junjun Kong, Xiangqun Chen
ISLPED4
2015 WuKong: a scalable and accurate two-phase approach to Android app clone detection
abstract
Repackaged Android applications (app clones) have been found in many third-party markets, which not only compromise the copyright of original authors, but also pose threats to security and privacy of mobile users. Both fine-grained and coarse-grained approaches have been proposed to detect app clones. However, fine-grained techniques employing complicated clone detection algorithms are difficult to scale to hundreds of thousands of apps, while coarse-grained techniques based on simple features are scalable but less accurate. This paper proposes WuKong, a two-phase detection approach that includes a coarse-grained detection phase to identify suspicious apps by comparing light-weight static semantic features, and a fine-grained phase to compare more detailed features for only those apps found in the first phase. To further improve the detection speed and accuracy, we also introduce an automated clustering-based preprocessing step to filter third-party libraries before conducting app clone detection. Experiments on more than 100,000 Android apps collected from five Android markets demonstrate the effectiveness and scalability of our approach.
Haoyu Wang 0001, Yao Guo 0001, Ziang Ma, Xiangqun Chen
ISSTA4
2015 SplitDroid: Isolated Execution of Sensitive Components for Mobile Applications
Yao Guo 0001, Xiangqun Chen
SecureComm3
2014 Preserving Location-Related Privacy Collaboratively in Geo-social Networks
abstract
The emerging geo-social networks bring us attractive location-based services as well as serious location-related privacy threats. Location information of users in geo-social networks might be revealed by friends carelessly, or deduced by users curiously or even maliciously. In order to avoid location leakages, we propose collaborative privacy management in geo-social networks. Users specify and broadcast their preferences on location-related privacies in advance, so that potential leakages can be reported automatically when new resources arrive. If necessary, the associated spatial and/or temporal information of resources will be tweaked according to the privacy preferences of involving users, so that "old" leakages can be eliminated while ensuring that "new" ones are not introduced. We design algorithms for such tweaks and construct experiments on a simulated dataset to demonstrate their usability and applicability.
Mingxuan Yuan, Yao Guo 0001, Xiangqun Chen, Lei Chen 0002
COMPSAC4
2014 An empirical study of indoor localization algorithms with densely deployed APs
abstract
Many indoor positioning algorithms have been proposed in the last decade, most of which are based on WiFi RSS fingerprints. However, the environment has changed dramatically since the original algorithms using only a few Access Points (APs). A typical building with densely deployed APs might contain hundreds of APs. The explosive growth of the number of APs introduces new challenges to these WiFi-based localization algorithms. This paper presents an empirical study of WiFi fingerprint-based indoor localization algorithms in a real-world environment with hundreds of APs. Our study aims to answer several important research questions regarding the influence of the number of APs, time variance and device variance. The study implements four existing algorithms and also proposes a new algorithm called LCS that is designed specifically for an AP-intensive environment. We compare the localization accuracy of different algorithms with different variances in the experimental results, which shows that the proposed LCS algorithm is able to efficiently resist diverse variances in an AP-intensive setup.
Xin Chen 0073, Junjun Kong, Yao Guo 0001, Xiangqun Chen
GLOBECOM4
2014 Topic Evolutions in Scientific Conferences
Yao Guo 0001, Xiangqun Chen, Weizhong Shao, Lei Chen 0002
SEKE3
2014 Similarity-based web browser optimization
abstract
The performance of web browsers has become a major bottleneck when dealing with complex webpages. Many calculation redundancies exist when processing similar webpages, thus it is possible to cache and reuse previously calculated intermediate results to improve web browser performance significantly. In this paper, we propose a similarity-based optimization approach to improve webpage processing performance of web browsers. Through caching and reusing of style properties calculated previously, we are able to eliminate the redundancies caused by processing similar webpages from the same website. We propose a tree-structured architecture to store style properties to facilitate efficient caching and reuse. Experiments on webpages of various websites show that the proposed technique can speed up the webpage loading process by up to 68% and reduce the redundant style calculations by up to 77% for the first visit to a webpage with almost negligible overhead.
Haoyu Wang 0001, Mengxin Liu, Yao Guo 0001, Xiangqun Chen
WWW4
2014 Context-aware usage control for web of things
abstract
ABSTRACT The Web of Things (WoT), inherited from the Internet of Things (IoT), encapsulates functionalities into publishable services on the Web to enable the IoT a seamless integration with the Web. The openness of the Web, in turn, directly exposes WoT to existing attacks from the Web. In addition, WoT possesses characteristics of high security and privacy concerns, mobility, and limited capabilities, which require specific and additional security and privacy protection beyond existing mechanisms. More importantly, WoT is inherently connected to its context, so context information must be taken into account in its security and privacy measures. To address these challenges, we propose a context‐aware usage control model (ConUCON), which leverages the context information to enhance data, resource, and service protection for WoT. On the basis of ConUCON, we also design and implement a context‐aware usage control framework on the middleware layer in our ongoing SmartHome project, to provide security and privacy protection. ConUCON is designed specifically to express the context‐aware usage policy specification, such that security and privacy requirements can be easily specified and enforced with the proposed model and framework. Finally, we apply ConUCON to a remote appliance management prototype, as a case study, to demonstrates its feasibility in a real environment. Copyright © 2012 John Wiley & Sons, Ltd.
Guangdong Bai, Liang Gu, Yao Guo 0001, Xiangqun Chen
Secur. Commun. Networks5
2013 Towards an operating system for the campus
abstract
Almost every computing device runs an operating system, which is responsible for managing different resources on the device and providing higher-level programming abstractions. This paper proposes CampusOS, an operating system which is responsible for managing networked resources on university campuses, including data of students, teachers, courses, organizations, and even data generated from users' computing devices. CampusOS provides flexible support for campus application development with SDKs consisting of campus-related APIs. CampusOS features and SDK APIs can also be extended by developers easily. We discuss the design of CampusOS, as well as its challenges.
Pengfei Yuan, Yao Guo 0001, Xiangqun Chen
Internetware3
2013 Power estimation for mobile applications with profile-driven battery traces
abstract
It becomes very important to understand power characteristics of mobile applications because more and more complex applications are running on modern smartphones. Although many techniques have been proposed to estimate the power dissipation rate for mobile applications, it typically requires hardware support (i.e., power meters) or complex power models (software profiling or hardware parameters). These techniques might work well in labs with a small set of applications. However, it becomes impractical when we try to estimate the power of mobile applications in an uncontrolled environment. This paper proposes a novel method for estimating the power consumption of mobile applications with profile-based battery traces. Battery traces can be easily collected through a user-level application on any devices. Although it is difficult to achieve accurate results for only a few users because battery changes are coarse-grained, the method is expected to reach an accurate estimation when the number of battery traces reaches a certain scale. Our experiments based on battery traces from more than 80,000 users demonstrate that it is possible to estimate application power with only coarse-grained battery traces. The results are also validated with measured power numbers from a Monsoon power monitor.
Chengke Wang, Fengrun Yan, Yao Guo 0001, Xiangqun Chen
ISLPED4
2012 Security model oriented attestation on dynamically reconfigurable component-based systems
Liang Gu, Guangdong Bai, Yao Guo 0001, Xiangqun Chen, Hong Mei 0001
J. Netw. Comput. Appl.4
2010 An Automatic Configuration Approach to Improve Real-Time Application Throughput While Attaining Determinism
abstract
Determinism and throughput are two important performance measures for Java-based real-time applications, but they often conflict. Therefore, it is significant to improve throughput for Java-based real-time applications while guaranteeing its execution time determinism. In this paper, we propose an automatic configuration approach to assign real-time thread priorities to solve the above-mentioned problem. In this approach, we propose an innovative representation of determinism related with real-time thread priorities using stochastic process. Java-based real-time application's throughput is quantified with thread priorities as parameters. The algorithm of integer programming is used to optimize throughput with boundary conditions of the level of determinism. Finally, the Sweet Factory application is tested to evaluate the effect of our approach. Experiment results show that throughput for Java-based real-time applications could be efficiently improved while keeping the execution time determinism with our approach.
Donggang Cao, Xiangqun Chen, Hong Mei 0001
COMPSAC4
2010 Context-Aware Usage Control for Android
Guangdong Bai, Liang Gu, Yao Guo 0001, Xiangqun Chen
SecureComm5
2009 FPValidator: Validating Type Equivalence of Function Pointers on the Fly
abstract
Validating function pointers dynamically is very useful for intrusion detection since many runtime attacks exploit function pointer vulnerabilities. Most current solutions tackle this problem through checking whether function pointers target the addresses within the code segment or, more strictly, valid function entries. However, they cannot detect function entry attacks that manipulate function pointers to target valid function entries but invoke them maliciously. This paper proposes FPValidator, a new solution capable of dynamically validating the type equivalence between function pointers and target functions, which can detect all function entry attacks that violate type equivalence. An effective and efficient type matching approach based on labeled type signature is proposed to perform fast type equivalence checking. The validation code and necessary type information are inserted by a compilation-stage instrumentation mechanism, bringing no extra burden to developers. We integrate FPValidator into GCC and evaluation shows that its performance overhead is only about 2%.
Yao Guo 0001, Xiangqun Chen
ACSAC3
2009 SAConf: Semantic Attestation of Software Configurations
Yao Guo 0001, Xiangqun Chen
ATC3
2009 Transaction-based adaptive dynamic voltage scaling for interactive applications
abstract
In an interactive embedded system, special task execution patterns and scheduling constraints exist due to frequent human-computer interactions. This paper proposes a transaction-based dynamic voltage scaling (T-DVS) approach that takes into account the characteristics of interactive transactions. T-DVS scales CPU performance levels to reduce energy consumption, while satisfying the constraints of both human-perceptual threshold and CPU requirement of an interactive transaction. T-DVS considers CPU requirements of both interactive and background tasks during a user interaction. It exploits CPU idle time waiting for user responses to run background task with lower CPU frequency. Experiments demonstrate that T-DVS can reduce energy consumption significantly compared to state-of-the-art approaches, with little sacrifice in user-perceived performance.
Yao Guo 0001, Xiangqun Chen
ISLPED3
2008 Keep Passwords Away from Memory: Password Caching and Verification Using TPM
abstract
TPM is able to provide strong secure storage for sensitive data such as passwords. Although several commercial password managers have used TPM to cache passwords, they are not capable of protecting passwords during verification. This paper proposes a new TPM-based password caching and verification method called PwdCaVe. In addition to using TPM in password caching, PwdCaVe also uses TPM during password verification. In PwdCaVe, all password-related computations are performed in the TPM. PwdCaVe guarantees that once a password is cached in the TPM, it will be protected by the TPM through the rest of its lifetime, thus eliminating the possibility that passwords might be attacked in memory. A prototype of PwdCaVe is implemented on Linux to demonstrate its feasibility.
Yao Guo 0001, Xiangqun Chen
AINA4
2008 Automated Aspect Recommendation through Clustering-Based Fan-in Analysis
abstract
Identifying code implementing a crosscutting concern (CCC) automatically can benefit the maintainability and evolvability of the application. Although many approaches have been proposed to identify potential aspects, a lot of manual work is typically required before these candidates can be converted into refactorable aspects. In this paper, we propose a new aspect mining approach, called clustering-based fan-in analysis (CBFA), to recommend aspect candidates in the form of method clusters, instead of single methods. CBFA uses a new lexical based clustering approach to identify method clusters and rank the clusters using a new ranking metric called cluster fan- in. Experiments on Linux and JHotDraw show that CBFA can provide accurate recommendations while improving aspect mining coverage significantly compared to other state-of-the-art mining approaches.
Danfeng Zhang, Yao Guo 0001, Xiangqun Chen
ASE3
2007 Toward Efficient Aspect Mining for Linux
abstract
Code implementing a crosscutting concern spreads over many parts of the Linux code. Identifying these code automatically can benefit both the maintainability and evolvability of Linux. In this paper, we present a case study on how to identify aspects in the Linux code. First, we analyze four typical crosscutting concerns in Linux and show how to apply existing mining approaches to identify these concerns. We then propose three new mining approaches and compare their performance with the original methods. Experiments show that the proposed mining approaches can find these concerns more efficiently in Linux.
Danfeng Zhang, Yao Guo 0001, Xiangqun Chen
APSEC4
2007 Towards a Software Framework for Building Highly Flexible Component-Based Embedded Operating Systems
Qiming Teng, Xiangqun Chen
EUC4
2005 A HAL for Component-Based Embedded Operating Systems
abstract
Many standards for operating system interfaces or on-chip bus interfaces have been developed. Component-based EOS projects are seeking approaches to adapt component based software development technologies to embedded systems. An important issue is to develop hardware independent software that meets differing requirements pertinent to embedded applications. The hardware abstraction layer (HAL) presented here is to serve this purpose. In JBEOS, a component based EOS developed at Peking Univ. The following aspects are stressed: abstraction of basic data types, including their in-memory representation and operations; encapsulation of H/W dependent features into clear interfaces for kernel and user applications; H/W peculiarities should be encapsulated but not masked, i.e., HAL should provide mechanisms to access H/W directly when desired; minimized ROM and RAM footprints; supports for EOS above with no bias to any specific design and/or implementation; minimum efforts required when porting to a new platform.
Qiming Teng, Xiangqun Chen
COMPSAC (2)3
2004 Extraction and Visualization of Architectural Structure Based on Cross References among Object Files
abstract
Reverse engineering of legacy systems is a knowledge-intensive process to reconstruct the understanding of a system. A semi-automatic process that can extract architecture level structure from legacy systems is introduced in This work. Exact facts related to cross-references among ELF objects are extracted from files automatically, and then partitioned into hierarchical groups by close cooperation between domain experts and an assistant tool DEREF. By resolving the cross references among these groups, the architectural structure is reconstructed and then visualized using auto-layout techniques. A case study on three embedded operating system demonstrates that this process can be used to obtain a comprehensive understanding about legacy systems even without any a priori knowledge about its design.
Qiming Teng, Xiangqun Chen, Lu Zhang 0023
COMPSAC2