VLDB 2026 Research / reviewers in the wild / expert
Xinyu Xing 0001
dblp:40/6431-1
· DBLP profile ↗
84ranked-venue papers
8as first author
35since 2021 · last 2025
0000-0001-6733-226XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 48 · 2 first-author · 23 since 2021Artificial intelligence and machine learning · 17 · 1 first-author · 10 since 2021Computer networks · 9 · 2 first-authorDatabases, data management, data science and information retrieval · 9 · 3 first-authorSoftware engineering, systems software and programming languages · 4 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-authorSystems, architecture and hardware · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | GPO: Learning from Critical Steps to Improve LLM ReasoningabstractLarge language models (LLMs) are increasingly used in various domains, showing impressive potential on various tasks.
Recently, reasoning LLMs have been proposed to improve the \textit{reasoning} or \textit{thinking} capabilities of LLMs to solve complex problems.
Despite the promising results of reasoning LLMs, enhancing the multi-step reasoning capabilities of LLMs still remains a significant challenge.
While existing optimization methods have advanced the LLM reasoning capabilities, they often treat reasoning trajectories as a whole, without considering the underlying critical steps within the trajectory. In this paper, we introduce \textbf{G}uided \textbf{P}ivotal \textbf{O}ptimization (GPO), a novel fine-tuning strategy that dives into the reasoning process to enable more effective improvements.
GPO first identifies the `critical step' within a reasoning trajectory - a point that the model must carefully proceed so as to succeed at the problem. We locate the critical step by estimating the advantage function.
GPO then resets the policy to the critical step and samples the new rollout and prioritizes learning process on those rollouts.
This focus allows the model to learn more effectively from pivotal moments within the reasoning process to improve the reasoning performance.
We demonstrate that GPO is not a standalone method, but rather a general strategy that can be integrated with various optimization methods to improve reasoning performance.
Besides theoretical analysis, our experiments across challenging reasoning benchmarks show that GPO can consistently and significantly enhances the performance of existing optimization methods, showcasing its effectiveness and generalizability in improving LLM reasoning by concentrating on pivotal moments within the generation process. Jiahao Yu 0001, Zelei Cheng, Xian Wu 0007, Xinyu Xing 0001 |
NeurIPS | 4 |
| 2025 | BlockScan: Detecting Anomalies in Blockchain TransactionsabstractWe propose BlockScan, a customized Transformer for anomaly detection in blockchain transactions.
Unlike existing methods that rely on rule-based systems or directly apply off-the-shelf large language models (LLMs), BlockScan introduces a series of customized designs to effectively model the unique data structure of blockchain transactions.
First, a blockchain transaction is multi-modal, containing blockchain-specific tokens, texts, and numbers.
We design a novel modularized tokenizer to handle these multi-modal inputs, balancing the information across different modalities.
Second, we design a customized masked language modeling mechanism for pretraining the Transformer architecture, incorporating RoPE embedding and FlashAttention for handling longer sequences.
Finally, we design a novel anomaly detection method based on the model outputs.
We further provide theoretical analysis for the detection method of our system.
Extensive evaluations on Ethereum and Solana transactions demonstrate BlockScan's exceptional capability in anomaly detection while maintaining a low false positive rate.
Remarkably, BlockScan is the only method that successfully detects anomalous transactions on Solana with high accuracy, whereas all other approaches achieved very low or zero detection recall scores.
This work sets a new benchmark for applying Transformer-based approaches in blockchain data analysis. Jiahao Yu 0001, Xian Wu 0007, Wenbo Guo 0002, Xinyu Xing 0001 |
NeurIPS | 5 |
| 2025 | Mind the Inconspicuous: Revealing the Hidden Weakness in Aligned LLMs' Refusal Boundaries
Jiahao Yu 0001, Haozheng Luo, Jerry Yao-Chieh Hu, Yan Chen 0004, Wenbo Guo 0002, Han Liu 0001, Xinyu Xing 0001 |
USENIX Security Symposium | 7 |
| 2025 | PATCHAGENT: A Practical Program Repair Agent Mimicking Human Expertise
Zheng Yu 0003, Yuhang Wu 0003, Jiahao Yu 0001, Meng Xu 0025, Dongliang Mu, Yan Chen 0004, Xinyu Xing 0001 |
USENIX Security Symposium | 8 |
| 2024 | TGRop: Top Gun of Return-Oriented Programming Automation
Nanyu Zhong, Yueqi Chen 0001, Yanyan Zou 0002, Xinyu Xing 0001, Jinwei Dong, Bingcheng Xian, Jiaxu Zhao 0004, Binghong Liu, Wei Huo 0005 |
ESORICS (3) | 4 |
| 2024 | RICE: Breaking Through the Training Bottlenecks of Reinforcement Learning with ExplanationabstractDeep reinforcement learning (DRL) is playing an increasingly important role in real-world applications. However, obtaining an optimally performing DRL agent for complex tasks, especially with sparse rewards, remains a significant challenge. The training of a DRL agent can be often trapped in a bottleneck without further progress. In this paper, we propose RICE, an innovative refining scheme for reinforcement learning that incorporates explanation methods to break through the training bottlenecks. The high-level idea of RICE is to construct a new initial state distribution that combines both the default initial states and critical states identified through explanation methods, thereby encouraging the agent to explore from the mixed initial states. Through careful design, we can theoretically guarantee that our refining scheme has a tighter sub-optimality bound. We evaluate RICE in various popular RL environments and real-world applications. The results demonstrate that RICE significantly outperforms existing refining schemes in enhancing agent performance. Zelei Cheng, Xian Wu 0007, Jiahao Yu 0001, Sabrina Yang, Gang Wang 0011, Xinyu Xing 0001 |
ICML | 6 |
| 2024 | Soft-Label Integration for Robust Toxicity ClassificationabstractToxicity classification in textual content remains a significant problem. Data with labels from a single annotator fall short of capturing the diversity of human perspectives. Therefore, there is a growing need to incorporate crowdsourced annotations for training an effective toxicity classifier. Additionally, the standard approach to training a classifier using empirical risk minimization (ERM) may fail to address the potential shifts between the training set and testing set due to exploiting spurious correlations. This work introduces a novel bi-level optimization framework that integrates crowdsourced annotations with the soft-labeling technique and optimizes the soft-label weights by Group Distributionally Robust Optimization (GroupDRO) to enhance the robustness against out-of-distribution (OOD) risk. We theoretically prove the convergence of our bi-level optimization algorithm. Experimental results demonstrate that our approach outperforms existing baseline methods in terms of both average and worst-group accuracy, confirming its effectiveness in leveraging crowdsourced annotations to achieve more effective and robust toxicity classification. Zelei Cheng, Xian Wu 0007, Jiahao Yu 0001, Xin-Qiang Cai, Xinyu Xing 0001 |
NeurIPS | 6 |
| 2024 | LLM-Fuzzer: Scaling Assessment of Large Language Model Jailbreaks
Jiahao Yu 0001, Xingwei Lin, Zheng Yu 0003, Xinyu Xing 0001 |
USENIX Security Symposium | 4 |
| 2024 | Take a Step Further: Understanding Page Spray in Linux Kernel Exploitation
Dang K. Le, Zhenpeng Lin, Kyle Zeng, Ruoyu Wang 0001, Tiffany Bao, Yan Shoshitaishvili, Adam Doupé, Xinyu Xing 0001 |
USENIX Security Symposium | 9 |
| 2024 | CAMP: Compiler and Allocator-based Heap Memory Protection
Zhenpeng Lin, Zheng Yu 0003, Simone Campanoni, Peter A. Dinda, Xinyu Xing 0001 |
USENIX Security Symposium | 6 |
| 2024 | SeaK: Rethinking the Design of a Secure Allocator for OS Kernel
Zicheng Wang 0010, Yicheng Guang, Yueqi Chen 0001, Zhenpeng Lin, Michael V. Le, Dang K. Le, Dan Williams 0001, Xinyu Xing 0001, Zhongshu Gu, Hani Jamjoom |
USENIX Security Symposium | 8 |
| 2024 | ShadowBound: Efficient Heap Memory Protection Through Advanced Metadata Management and Customized Compiler Optimization
Zheng Yu 0003, Ganxiang Yang, Xinyu Xing 0001 |
USENIX Security Symposium | 3 |
| 2024 | Towards Unveiling Exploitation Potential With Multiple Error Behaviors for Kernel BugsabstractNowadays, fuzz testing has significantly expedited the vulnerability discovery of Linux kernel. Security analysts use the manifested error behaviors to infer the exploitability of one bug and thus prioritize the patch development. However, only using an error behavior in the report, security analysts might underestimate the exploitability of the kernel bug because it could manifest various error behaviors indicating different exploitation potentials. In this work, we conduct an empirical study on multiple error behaviors of kernel bugs to understand 1) the prevalence of multiple error behaviors and the possible impact of multiple error behaviors towards the exploitation potential; 2) the factors that manifest multiple error behaviors with different exploitation potential. We collectedall the fixed kernel bugsreported on Syzbot from September 2017 to January 2022, including 3,352 bug reports. We observed that multiple error behaviors manifested by kernel bugs are prevalent in the real world, and more error behaviors help unveil the exploitability of kernel bugs. Then we organized Linux kernel experts to analyze a sample of kernel bug dataset (484 bug reports, unique 162 bugs) and identified 6 key contributing factors to the mutiple error behaviors. Finally, based on the empirical findings, we propose an object-driven fuzzing technique to explore all possible error behaviors that a kernel bug might bring about. To evaluate the utility of our proposed technique, we implement our fuzzing toolGREBEand apply it to 60 real-world Linux kernel bugs. On average,GREBEcould manifest 2+ additional error behaviors for each of the kernel bugs. For 26 kernel bugs,GREBEdiscovers higher exploitation potential. We report to kernel vendors some of the bugs – the exploitability of which was wrongly assessed and the corresponding patch has not yet been carefully applied – resulting in their rapid patch adoption. Ziqin Liu, Zhenpeng Lin, Yueqi Chen 0001, Yuhang Wu 0003, Yalong Zou, Dongliang Mu, Xinyu Xing 0001 |
IEEE Trans. Dependable Secur. Comput. | 7 |
| 2023 | TGC: Transaction Graph Contrast Network for Ethereum Phishing Scam DetectionabstractPhishing scams have become the most serious type of crime involved in Ethereum. However, existing methods ignore the natural camouflage and sparse distribution of phishing scams in Ethereum leading to unsatisfactory performance, and they are also limited by the data scale which cannot be applied to real-world dynamic scenarios. In this paper, we propose a Transaction Graph Contrast network (TGC) to enhance phishing scam detection performance on Ethereum. TGC inputs subgraphs instead of the entire graph for training, which eases the model’s requirements for machine configuration and data connectivity. Motivated by phishing nodes are surrounded by normal nodes, we design the comparison between node-level to help phishing nodes learn the unique properties of themselves different from their neighbors. Observing the small number and sparse distribution of phishing nodes, we narrow the distance between phishing nodes by comparing node context-level structures, so as to learn universal transaction patterns. We further combine the obtained features with common statistics to identify phishing addresses. Evaluated on real-world Ethereum phishing scams datasets, our TGC outperforms the state-of-the-art methods in detecting phishing addresses and has obvious advantages in large-scale and dynamic scenarios. Gaopeng Gou, Chang Liu 0049, Gang Xiong 0001, Zhen Li 0011, Junchao Xiao, Xinyu Xing 0001 |
ACSAC | 7 |
| 2023 | RetSpill: Igniting User-Controlled Data to Burn Away Linux Kernel ProtectionsabstractLeveraging a control flow hijacking primitive (CFHP) to gain root privileges is critical to attackers striving to exploit Linux kernel vulnerabilities. Such attack has become increasingly elusive as security researchers propose capable kernel security mitigations, leading to the development of complex (and, as a trade-off, brittle and unreliable) attack techniques to regain it. In this paper, we obviate the need for complexity by proposing RetSpill, a powerful yet elegant exploitation technique that employs user space data already present on the kernel stack for privilege escalation. Kyle Zeng, Zhenpeng Lin, Kangjie Lu, Xinyu Xing 0001, Ruoyu Wang 0001, Adam Doupé, Yan Shoshitaishvili, Tiffany Bao |
CCS | 4 |
| 2023 | StateMask: Explaining Deep Reinforcement Learning through State MaskabstractDespite the promising performance of deep reinforcement learning (DRL) agents in many challenging scenarios, the black-box nature of these agents greatly limits their applications in critical domains. Prior research has proposed several explanation techniques to understand the deep learning-based policies in RL. Most existing methods explain why an agent takes individual actions rather than pinpointing the critical steps to its final reward. To fill this gap, we propose StateMask, a novel method to identify the states most critical to the agent's final reward. The high-level idea of StateMask is to learn a mask net that blinds a target agent and forces it to take random actions at some steps without compromising the agent's performance. Through careful design, we can theoretically ensure that the masked agent performs similarly to the original agent. We evaluate StateMask in various popular RL environments and show its superiority over existing explainers in explanation fidelity. We also show that StateMask has better utilities, such as launching adversarial attacks and patching policy errors. Zelei Cheng, Xian Wu 0007, Jiahao Yu 0001, Wenhai Sun, Wenbo Guo 0002, Xinyu Xing 0001 |
NeurIPS | 6 |
| 2023 | From Grim Reality to Practical Solution: Malware Classification in Real-World NoiseabstractMalware datasets inevitably contain incorrect labels due to the shortage of expertise and experience needed for sample labeling. Previous research demonstrated that a training dataset with incorrectly labeled samples would result in inaccurate model learning. To address this problem, researchers have proposed various noise learning methods to offset the impact of incorrectly labeled samples, and in image recognition and text mining applications, these methods demonstrated great success. In this work, we apply both representative and state-of-the-art noise learning methods to real-world malware classification tasks. We surprisingly observe that none of the existing methods could minimize incorrect labels’ impact. Through a carefully designed experiment, we discover that the inefficacy mainly results from extreme data imbalance and the high percentage of incorrectly labeled data samples. As such, we further propose a new noise learning method and name it after MORSE. Unlike existing methods, MORSE customizes and extends a state-of-the-art semi-supervised learning technique. It takes possibly incorrectly labeled data as unlabeled data and thus avoids their potential negative impact on model learning. In MORSE, we also integrate a sample re-weighting method that balances the training data usage in the model learning and thus handles the data imbalance challenge. We evaluate MORSE on both our synthesized and real-world datasets. We show that MORSE could significantly outperform existing noise learning methods and minimize the impact of incorrectly labeled data. Xian Wu 0007, Wenbo Guo 0002, Baris Coskun, Xinyu Xing 0001 |
SP | 5 |
| 2023 | PATROL: Provable Defense against Adversarial Policy in Two-player Games
Wenbo Guo 0002, Xian Wu 0007, Lun Wang 0001, Xinyu Xing 0001, Dawn Song |
USENIX Security Symposium | 4 |
| 2023 | Mitigating Security Risks in Linux with KLAUS: A Method for Evaluating Patch Correctness
Yuhang Wu 0003, Zhenpeng Lin, Yueqi Chen 0001, Dang K. Le, Dongliang Mu, Xinyu Xing 0001 |
USENIX Security Symposium | 6 |
| 2023 | AIRS: Explanation for Deep Reinforcement Learning based Security Applications
Jiahao Yu 0001, Wenbo Guo 0002, Gang Wang 0011, Ting Wang 0006, Xinyu Xing 0001 |
USENIX Security Symposium | 6 |
| 2022 | DirtyCred: Escalating Privilege in Linux KernelabstractThe kernel vulnerability DirtyPipe was reported to be present in nearly all versions of Linux since 5.8. Using this vulnerability, a bad actor could fulfill privilege escalation without triggering existing kernel protection and exploit mitigation, making this vulnerability particularly disconcerting. However, the success of DirtyPipe exploitation heavily relies on this vulnerability's capability (i.e., injecting data into the arbitrary file through Linux's pipes). Such an ability is rarely seen for other kernel vulnerabilities, making the defense relatively easy. As long as Linux users eliminate the vulnerability, the system could be relatively secure. Zhenpeng Lin, Yuhang Wu 0003, Xinyu Xing 0001 |
CCS | 3 |
| 2022 | An In-depth Analysis of Duplicated Linux Kernel Bug Reports
Dongliang Mu, Yuhang Wu 0003, Yueqi Chen 0001, Zhenpeng Lin, Chensheng Yu, Xinyu Xing 0001, Gang Wang 0011 |
NDSS | 6 |
| 2022 | DeJITLeak: eliminating JIT-induced timing side-channel leaksabstractTiming side-channels can be exploited to infer secret information when the execution time of a program is correlated with secrets. Recent work has shown that Just-In-Time (JIT) compilation can introduce new timing side-channels in programs even if they are time-balanced at the source code level. In this paper, we propose a novel approach to eliminate JIT-induced leaks. We first formalise timing side-channel security under JIT compilation via the notion of time-balancing, laying the foundation for reasoning about programs with JIT compilation. We then propose to eliminate JIT-induced leaks via a fine-grained JIT compilation. To this end, we provide an automated approach to generate compilation policies and a novel type system to guarantee its soundness. We develop a tool DeJITLeak for real-world Java and implement the fine-grained JIT compilation in HotSpot JVM. Experimental results show that DeJITLeak can effectively and efficiently eliminate JIT-induced leaks on three widely adopted benchmarks in the setting of side-channel detection. JulianAndres JiYang, Fu Song, Taolue Chen 0001, Xinyu Xing 0001 |
ESEC/SIGSOFT FSE | 5 |
| 2022 | GREBE: Unveiling Exploitation Potential for Linux Kernel BugsabstractNowadays, dynamic testing tools have significantly expedited the discovery of bugs in the Linux kernel. When unveiling kernel bugs, they automatically generate reports, specifying the errors the Linux encounters. The error in the report implies the possible exploitability of the corresponding kernel bug. As a result, many security analysts use the manifested error to infer a bug’s exploitability and thus prioritize their exploit development effort. However, using the error in the report, security researchers might underestimate a bug’s exploitability. The error exhibited in the report may depend upon how the bug is triggered. Through different paths or under different contexts, a bug may manifest various error behaviors implying very different exploitation potentials. This work proposes a new kernel fuzzing technique to explore all the possible error behaviors that a kernel bug might bring about. Unlike conventional kernel fuzzing techniques concentrating on kernel code coverage, our fuzzing technique is more directed towards the buggy code fragment. It introduces an object-driven kernel fuzzing technique to explore various contexts and paths to trigger the reported bug, making the bug manifest various error behaviors. With the newly demonstrated errors, security researchers could better infer a bug’s possible exploitability. To evaluate our proposed technique’s effectiveness, efficiency, and impact, we implement our fuzzing technique as a tool GREBE and apply it to 60 real-world Linux kernel bugs. On average, GREBE could manifest 2+ additional error behaviors for each of the kernel bugs. For 26 kernel bugs, GREBE discovers higher exploitation potential. We report to kernel vendors some of the bugs – the exploitability of which was wrongly assessed and the corresponding patch has not yet been carefully applied – resulting in their rapid patch adoption. Zhenpeng Lin, Yueqi Chen 0001, Yuhang Wu 0003, Dongliang Mu, Chensheng Yu, Xinyu Xing 0001 |
SP | 6 |
| 2022 | Playing for K(H)eaps: Understanding and Improving Linux Kernel Exploit Reliability
Kyle Zeng, Yueqi Chen 0001, Haehyun Cho, Xinyu Xing 0001, Adam Doupé, Yan Shoshitaishvili, Tiffany Bao |
USENIX Security Symposium | 4 |
| 2021 | Facilitating Vulnerability Assessment through PoC MigrationabstractRecent research shows that, even for vulnerability reports archived by MITRE/NIST, they usually contain incomplete information about the software's vulnerable versions, making users of under-reported vulnerable versions at risk. In this work, we address this problem by introducing a fuzzing-based method. Technically, this approach first collects the crashing trace on the reference version of the software. Then, it utilizes the trace to guide the mutation of the PoC input so that the target version could follow the trace similar to the one observed on the reference version. Under the mutated input, we argue that the target version's execution could have a higher chance of triggering the bug and demonstrating the vulnerability's existence. We implement this idea as an automated tool, named VulScope. Using 30 real-world CVEs on 470 versions of software, VulScope is demonstrated to introduce no false positives and only 7.9% false negatives while migrating PoC from one version to another. Besides, we also compare our method with two representative fuzzing tools AFL and AFLGO. We find VulScope outperforms both of these existing techniques while taking the task of PoC migration. Finally, by using VulScope, we identify 330 versions of software that MITRE/NIST fails to report as vulnerable. Jiarun Dai, Yuan Zhang 0009, Hailong Xu, Haiming Lyu, Zicheng Wu, Xinyu Xing 0001, Min Yang 0002 |
CCS | 6 |
| 2021 | Adversarial Policy Learning in Two-player Competitive GamesabstractIn a two-player deep reinforcement learning task, recent work shows an attacker could learn an adversarial policy that triggers a target agent to perform poorly and even react in an undesired way. However, its efficacy heavily relies upon the zero-sum assumption made in the two-player game. In this work, we propose a new adversarial learning algorithm. It addresses the problem by resetting the optimization goal in the learning process and designing a new surrogate optimization function. Our experiments show that our method significantly improves adversarial agents’ exploitability compared with the state-of-art attack. Besides, we also discover that our method could augment an agent with the ability to abuse the target game’s unfairness. Finally, we show that agents adversarially re-trained against our adversarial agents could obtain stronger adversary-resistance. Wenbo Guo 0002, Xian Wu 0007, Sui Huang, Xinyu Xing 0001 |
ICML | 4 |
| 2021 | DANCE: Enhancing saliency maps using decoysabstractSaliency methods can make deep neural network predictions more interpretable by identifying a set of critical features in an input sample, such as pixels that contribute most strongly to a prediction made by an image classifier. Unfortunately, recent evidence suggests that many saliency methods poorly perform, especially in situations where gradients are saturated, inputs contain adversarial perturbations, or predictions rely upon inter-feature dependence. To address these issues, we propose a framework, DANCE, which improves the robustness of saliency methods by following a two-step procedure. First, we introduce a perturbation mechanism that subtly varies the input sample without changing its intermediate representations. Using this approach, we can gather a corpus of perturbed ("decoy") data samples while ensuring that the perturbed and original input samples follow similar distributions. Second, we compute saliency maps for the decoy samples and propose a new method to aggregate saliency maps. With this design, we offset influence of gradient saturation. From a theoretical perspective, we show that the aggregated saliency map not only captures inter-feature dependence but, more importantly, is robust against previously described adversarial perturbation methods. Our empirical results suggest that, both qualitatively and quantitatively, DANCE outperforms existing methods in a variety of application domains. Yang Young Lu, Wenbo Guo 0002, Xinyu Xing 0001, William Stafford Noble |
ICML | 3 |
| 2021 | RNNRepair: Automatic RNN Repair via Model-based AnalysisabstractDeep neural networks are vulnerable to adversarial attacks. Due to their black-box nature, it is rather challenging to interpret and properly repair these incorrect behaviors. This paper focuses on interpreting and repairing the incorrect behaviors of Recurrent Neural Networks (RNNs). We propose a lightweight model-based approach (RNNRepair) to help understand and repair incorrect behaviors of an RNN. Specifically, we build an influence model to characterize the stateful and statistical behaviors of an RNN over all the training data and to perform the influence analysis for the errors. Compared with the existing techniques on influence function, our method can efficiently estimate the influence of existing or newly added training samples for a given prediction at both sample level and segmentation level. Our empirical evaluation shows that the proposed influence model is able to extract accurate and understandable features. Based on the influence model, our proposed technique could effectively infer the influential instances from not only an entire testing sequence but also a segment within that sequence. Moreover, with the sample-level and segment-level influence relations, RNNRepair could further remediate two types of incorrect predictions at the sample level and segment level. Xiaofei Xie, Wenbo Guo 0002, Lei Ma 0003, Wei Le, Jian Wang 0067, Lingjun Zhou, Yang Liu 0003, Xinyu Xing 0001 |
ICML | 8 |
| 2021 | BACKDOORL: Backdoor Attack against Competitive Reinforcement LearningabstractRecent research has confirmed the feasibility of backdoor attacks in deep reinforcement learning (RL) systems. However, the existing attacks require the ability to arbitrarily modify an agent's observation, constraining the application scope to simple RL systems such as Atari games. In this paper, we migrate backdoor attacks to more complex RL systems involving multiple agents and explore the possibility of triggering the backdoor without directly manipulating the agent's observation. As a proof of concept, we demonstrate that an adversary agent can trigger the backdoor of the victim agent with its own action in two-player competitive RL systems. We prototype and evaluate BackdooRL in four competitive environments. The results show that when the backdoor is activated, the winning rate of the victim drops by 17% to 37% compared to when not activated. The videos are hosted at https://github.com/wanglun1996/multi_agent_rl_backdoor_videos. Lun Wang 0001, Zaynah Javed, Xian Wu 0007, Wenbo Guo 0002, Xinyu Xing 0001, Dawn Song |
IJCAI | 5 |
| 2021 | FARE: Enabling Fine-grained Attack Categorization under Low-quality Labeled Data
Wenbo Guo 0002, Tongbo Luo, Vasant G. Honavar, Gang Wang 0011, Xinyu Xing 0001 |
NDSS | 6 |
| 2021 | EDGE: Explaining Deep Reinforcement Learning PoliciesabstractWith the rapid development of deep reinforcement learning (DRL) techniques, there is an increasing need to understand and interpret DRL policies. While recent research has developed explanation methods to interpret how an agent determines its moves, they cannot capture the importance of actions/states to a game's final result. In this work, we propose a novel self-explainable model that augments a Gaussian process with a customized kernel function and an interpretable predictor. Together with the proposed model, we also develop a parameter learning procedure that leverages inducing points and variational inference to improve learning efficiency. Using our proposed model, we can predict an agent's final rewards from its game episodes and extract time step importance within episodes as strategy-level explanations for that agent. Through experiments on Atari and MuJoCo games, we verify the explanation fidelity of our method and demonstrate how to employ interpretation to understand agent behavior, discover policy vulnerabilities, remediate policy errors, and even defend against adversarial attacks. Wenbo Guo 0002, Xian Wu 0007, Usmann Khan, Xinyu Xing 0001 |
NeurIPS | 4 |
| 2021 | Adversarial Policy Training against Deep Reinforcement Learning
Xian Wu 0007, Wenbo Guo 0002, Hua Wei 0001, Xinyu Xing 0001 |
USENIX Security Symposium | 4 |
| 2021 | CADE: Detecting and Explaining Concept Drift Samples for Security Applications
Wenbo Guo 0002, Qingying Hao, Arridhana Ciptadi, Ali Ahmadzadeh, Xinyu Xing 0001, Gang Wang 0011 |
USENIX Security Symposium | 6 |
| 2021 | POMP++: Facilitating Postmortem Program Diagnosis with Value-Set AnalysisabstractWith the emergence of hardware-assisted processor tracing, execution traces can be logged with lower runtime overhead and integrated into the core dump. In comparison with an ordinary core dump, such a new post-crash artifact provides software developers and security analysts with more clues to a program crash. However, existing works only rely on the resolved runtime information, which leads to the limitation in data flow recovery within long execution traces. In this work, we propose POMP++, an automated tool to facilitate the analysis of post-crash artifacts. More specifically, POMP++ introduces a reverse execution mechanism to construct the data flow that a program followed prior to its crash. Furthermore, POMP++ utilizes Value-set Analysis, which helps to verify memory alias relation, to improve the ability of data flow recovery. With the restored data flow, POMP++ then performs backward taint analysis and highlights program statements that actually contribute to the crash. We have implemented POMP++ for Linux system on x86-32 platform, and tested it against various crashes resulting from 31 distinct real-world security vulnerabilities. The evaluation shows that, our work can pinpoint the root causes in 29 cases, increase the number of recovered memory addresses by 12 percent and reduce the execution time by 60 percent compared with existing reverse execution. In short, POMP++ can accurately and efficiently pinpoint program statements that truly contribute to the crashes, making failure diagnosis significantly convenient. Dongliang Mu, Yunlan Du, Jianhao Xu, Jun Xu 0024, Xinyu Xing 0001, Bing Mao 0001, Peng Liu 0005 |
IEEE Trans. Software Eng. | 5 |
| 2020 | A Systematic Study of Elastic Objects in Kernel ExploitationabstractRecent research has proposed various methods to perform kernel exploitation and bypass kernel protection. For example, security researchers have demonstrated an exploitation method that utilizes the characteristic of elastic kernel objects to bypass KASLR, disclose stack/heap cookies, and even perform arbitrary read in the kernel. While this exploitation method is considered a commonly adopted approach to disclosing critical kernel information, there is no evidence indicating a strong need for developing a new defense mechanism to limit this exploitation method. It is because the effectiveness of this exploitation method is demonstrated only on anecdotal kernel vulnerabilities. It is unclear whether such a method is useful for a majority of kernel vulnerabilities. Yueqi Chen 0001, Zhenpeng Lin, Xinyu Xing 0001 |
CCS | 3 |
| 2020 | PDiff: Semantic-based Patch Presence Testing for Downstream KernelsabstractOpen-source kernels have been adopted by massive downstream vendors on billions of devices. However, these vendors often omit or delay the adoption of patches released in the mainstream version. Even worse, many vendors are not publicizing the patching progress or even disclosing misleading information. However, patching status is critical for groups (e.g., governments and enterprise users) that are keen to security threats. Such a practice motivates the need for reliable patch presence testing for downstream kernels. Currently, the best means of patch presence testing is to examine the existence of a patch in the target kernel by using the code signature match. However, such an approach cannot address the key challenges in practice. Specifically, downstream vendors widely customize the mainstream code and use non-standard building configurations, which often change the code around the patching sites such that the code signatures are ineffective. Zheyue Jiang, Yuan Zhang 0009, Jun Xu 0024, Zhenghe Wang, Xiaohan Zhang 0001, Xinyu Xing 0001, Min Yang 0002, Zhemin Yang |
CCS | 7 |
| 2020 | HART: Hardware-Assisted Kernel Module Tracing on Arm
Yunlan Du, Zhenyu Ning, Jun Xu 0024, Yueh-Hsun Lin, Fengwei Zhang, Xinyu Xing 0001, Bing Mao 0001 |
ESORICS (1) | 7 |
| 2020 | Towards Inspecting and Eliminating Trojan Backdoors in Deep Neural NetworksabstractA trojan backdoor is a hidden pattern typically implanted in a deep neural network (DNN). It could be activated and thus forces that infected model to behave abnormally when an input sample with a particular trigger is fed to that model. As such, given a DNN and clean input samples, it is challenging to inspect and determine the existence of a trojan backdoor. Recently, researchers design and develop several pioneering solutions to address this problem. They demonstrate that the proposed techniques have great potential in trojan detection. However, we show that none of these existing techniques completely address the problem. On the one hand, they mostly work under an unrealistic assumption of assuming the availability of the contaminated training database. On the other hand, these techniques can neither accurately detect the existence of trojan backdoors, nor restore high-fidelity triggers, especially when infected models are trained with high-dimensional data, and the triggers pertaining to the trojan vary in size, shape, and position. In this work, we propose TABOR, a new trojan detection technique. Conceptually, it formalizes the detection of a trojan backdoor as solving an optimization objective function. Different from the existing technique which also models trojan detection as an optimization problem, TABOR first designs a new objective function that could guide optimization to identify a trojan backdoor more correctly and accurately. Second, TABOR borrows the idea of interpretable AI to further prune the restored triggers. Last, TABOR designs a new anomaly detection method, which could not only facilitate the identification of intentionally injected triggers but also filter out false alarms (i.e., triggers detected from an uninfected model). We train 112 DNNs on five datasets and infect these models with two existing trojan attacks. We evaluate TABOR by using these infected models, and demonstrate that TABOR has much better performance in trigger restoration, trojan detection, and elimination than Neural Cleanse, the state-of-the-art trojan detection technique. Wenbo Guo 0002, Lun Wang 0001, Yan Xu 0019, Xinyu Xing 0001, Min Du 0003, Dawn Song |
ICDM | 4 |
| 2020 | BScout: Direct Whole Patch Presence Test for Java Executables
Jiarun Dai, Yuan Zhang 0009, Zheyue Jiang, Yingtian Zhou, Xinyu Xing 0001, Xiaohan Zhang 0001, Min Yang 0002, Zhemin Yang |
USENIX Security Symposium | 6 |
| 2020 | Tainting-Assisted and Context-Migrated Symbolic Execution of Android Framework for Vulnerability Discovery and Exploit GenerationabstractAndroid Application Framework is an integral and foundational part of the Android system. Each of the two billion (as of 2017) Android devices relies on the system services of Android Framework to manage applications and system resources. Given its critical role, a vulnerability in the framework can be exploited to launch large-scale cyber attacks and cause severe harms to user security and privacy. Recently, many vulnerabilities in Android Framework were exposed, showing that it is indeed vulnerable and exploitable. While there is a large body of studies on Android application analysis, research on Android Framework analysis is very limited. In particular, to our knowledge, there is no prior work that investigates how to enable symbolic execution of the framework, an approach that has proven to be very powerful for vulnerability discovery and exploit generation. We design and build the first system, Centaur, that enables symbolic execution of Android Framework. Due to the middleware nature and technical peculiarities of the framework that impinge on the analysis, many unique challenges arise and are addressed in Centaur. The system has been applied to discovering new vulnerability instances, which can be exploited by recently uncovered attacks against the framework, and to generating PoC exploits. Lannan Luo, Qiang Zeng 0001, Chen Cao 0004, Kai Chen 0012, Jian Liu 0008, Neng Gao, Min Yang 0002, Xinyu Xing 0001, Peng Liu 0005 |
IEEE Trans. Mob. Comput. | 9 |
| 2019 | PTrix: Efficient Hardware-Assisted Fuzzing for COTS BinaryabstractDespite its effectiveness in uncovering software defects, American Fuzzy Lop (AFL), one of the best grey-box fuzzers, is inefficient when fuzz-testing source-unavailable programs. AFL's binary-only fuzzing mode, QEMU-AFL, is typically 2-5× slower than its source- available fuzzing mode. The slowdown is largely caused by the heavy dynamic instrumentation. Recent fuzzing techniques use Intel Processor Tracing (PT), a light-weight tracing feature supported by recent Intel CPUs, to re- move the need of dynamic instrumentation. However, we found that these PT-based fuzzing techniques are even slower than QEMU-AFL when fuzzing real-world programs, making them less effective than QEMU-AFL. This poor performance is caused by the slow extraction of code coverage information from highly compressed PT traces. In this work, we present the design and implementation of PTrix, which fully unleashes the benefits of PT for fuzzing via three novel techniques. First, PTrix introduces a scheme to highly parallel the processing of PT trace and target program execution. Second, it directly takes decoded PT trace as feedback for fuzzing, avoiding the expensive reconstruction of code coverage information. Third, PTrix maintains the new feedback with stronger feedback than edge-based code coverage, which helps reach new code space and defects that AFL may not. We evaluated PTrix by comparing its performance with the state- of-the-art fuzzers. Our results show that, given the same amount of time, PTrix achieves a significantly higher fuzzing speed and reaches into code regions missed by the other fuzzers. In addition, PTrix identifies 35 new vulnerabilities in a set of previously well- fuzzed binaries, showing its ability to complement existing fuzzers. Yaohui Chen 0001, Dongliang Mu, Jun Xu 0024, Zhichuang Sun, Wenbo Shen, Xinyu Xing 0001, Long Lu, Bing Mao 0001 |
AsiaCCS | 6 |
| 2019 | SLAKE: Facilitating Slab Manipulation for Exploiting Vulnerabilities in the Linux KernelabstractTo determine the exploitability for a kernel vulnerability, a secu- rity analyst usually has to manipulate slab and thus demonstrate the capability of obtaining the control over a program counter or performing privilege escalation. However, this is a lengthy process because (1) an analyst typically has no clue about what objects and system calls are useful for kernel exploitation and (2) he lacks the knowledge of manipulating a slab and obtaining the desired layout. In the past, researchers have proposed various techniques to facilitate exploit development. Unfortunately, none of them can be easily applied to address these challenges. On the one hand, this is because of the complexity of the Linux kernel. On the other hand, this is due to the dynamics and non-deterministic of slab variations. In this work, we tackle the challenges above from two perspectives. First, we use static and dynamic analysis techniques to explore the kernel objects, and the corresponding system calls useful for exploitation. Second, we model commonly-adopted exploitation methods and develop a technical approach to facilitate the slab layout adjustment. By extending LLVM as well as Syzkaller, we implement our techniques and name their combination after SLAKE. We evaluate SLAKE by using 27 real-world kernel vulnerabilities, demonstrating that it could not only diversify the ways to perform kernel exploitation but also sometimes escalate the exploitability of kernel vulnerabilities. Yueqi Chen 0001, Xinyu Xing 0001 |
CCS | 2 |
| 2019 | Log2vec: A Heterogeneous Graph Embedding Based Approach for Detecting Cyber Threats within EnterpriseabstractConventional attacks of insider employees and emerging APT are both major threats for the organizational information system. Existing detections mainly concentrate on users' behavior and usually analyze logs recording their operations in an information system. In general, most of these methods consider sequential relationship among log entries and model users' sequential behavior. However, they ignore other relationships, inevitably leading to an unsatisfactory performance on various attack scenarios. We propose log2vec, a heterogeneous graph embedding based modularized method. First, it involves a heuristic approach that converts log entries into a heterogeneous graph in the light of diverse relationships among them. Next, it utilizes an improved graph embedding appropriate to the above heterogeneous graph, which can automatically represent each log entry into a low-dimension vector. The third component of log2vec is a practical detection algorithm capable of separating malicious and benign log entries into different clusters and identifying malicious ones. We implement a prototype of log2vec. Our evaluation demonstrates that log2vec remarkably outperforms state-of-the-art approaches, such as deep learning and hidden markov model (HMM). Besides, log2vec shows its capability to detect malicious events in various attack scenarios. Fucheng Liu, Yu Wen 0001, Dongxue Zhang, Xihe Jiang, Xinyu Xing 0001, Dan Meng 0002 |
CCS | 5 |
| 2019 | Errors, Misunderstandings, and Attacks: Analyzing the Crowdsourcing Process of Ad-blocking SystemsabstractAd-blocking systems such as Adblock Plus rely on crowdsourcing to build and maintain filter lists, which are the basis for determining which ads to block on web pages. In this work, we seek to advance our understanding of the ad-blocking community as well as the errors and pitfalls of the crowdsourcing process. To do so, we collected and analyzed a longitudinal dataset that covered the dynamic changes of popular filter-list EasyList for nine years and the error reports submitted by the crowd in the same period. Mshabab Alrizah, Sencun Zhu, Xinyu Xing 0001, Gang Wang 0011 |
Internet Measurement Conference | 3 |
| 2019 | RENN: Efficient Reverse Execution with Neural-Network-Assisted Alias AnalysisabstractReverse execution and coredump analysis have long been used to diagnose the root cause of software crashes. Each of these techniques, however, face inherent challenges, such as insufficient capability when handling memory aliases. Recent works have used hypothesis testing to address this drawback, albeit with high computational complexity, making them impractical for real world applications. To address this issue, we propose a new deep neural architecture, which could significantly improve memory alias resolution. At the high level, our approach employs a recurrent neural network (RNN) to learn the binary code pattern pertaining to memory accesses. It then infers the memory region accessed by memory references. Since memory references to different regions naturally indicate a non-alias relationship, our neural architecture can greatly reduce the burden of doing hypothesis testing to track down non-alias relation in binary code. Different from previous researches that have utilized deep learning for other binary analysis tasks, the neural network proposed in this work is fundamentally novel. Instead of simply using off-the-shelf neural networks, we designed a new recurrent neural architecture that could capture the data dependency between machine code segments. To demonstrate the utility of our deep neural architecture, we implement it as RENN, a neural network-assisted reverse execution system. We utilize this tool to analyze software crashes corresponding to 40 memory corruption vulnerabilities from the real world. Our experiments show that RENN can significantly improve the efficiency of locating the root cause for the crashes. Compared to a state-of-the-art technique, RENN has 36.25% faster execution time on average, detects an average of 21.35% more non-alias pairs, and successfully identified the root cause of 12.5% more cases. Dongliang Mu, Wenbo Guo 0002, Alejandro Cuevas, Yueqi Chen 0001, Jinxuan Gai, Xinyu Xing 0001, Bing Mao 0001, Chengyu Song |
ASE | 6 |
| 2019 | Towards the Detection of Inconsistencies in Public Security Vulnerability Reports
Wenbo Guo 0002, Yueqi Chen 0001, Xinyu Xing 0001, Gang Wang 0011 |
USENIX Security Symposium | 4 |
| 2019 | DEEPVSA: Facilitating Value-set Analysis with Deep Learning for Postmortem Program Analysis
Wenbo Guo 0002, Dongliang Mu, Xinyu Xing 0001, Min Du 0003, Dawn Song |
USENIX Security Symposium | 3 |
| 2019 | KEPLER: Facilitating Control-flow Hijacking Primitive Evaluation for Linux Kernel Vulnerabilities
Wei Wu 0010, Yueqi Chen 0001, Xinyu Xing 0001 |
USENIX Security Symposium | 3 |
| 2019 | All Your Clicks Belong to Me: Investigating Click Interception on the Web
Mingxue Zhang 0001, Wei Meng 0001, Sangho Lee 0001, Byoungyoung Lee, Xinyu Xing 0001 |
USENIX Security Symposium | 5 |
| 2019 | From proof-of-concept to exploitableabstractExploitability assessment of vulnerabilities is important for both defenders and attackers. The ultimate way to assess the exploitability is crafting a working exploit. However, it usually takes tremendous hours and significant manual efforts. To address this issue, automated techniques can be adopted. Existing solutions usually explore in depth the crashing paths , i.e., paths taken by proof-of-concept (PoC) inputs triggering vulnerabilities, and assess exploitability by finding exploitable states along the paths. However, exploitable states do not always exist in crashing paths. Moreover, existing solutions heavily rely on symbolic execution and are not scalable in path exploration and exploit generation. In this paper, we propose a novel solution to generate exploit for userspace programs or facilitate the process of crafting a kernel UAF exploit. Technically, we utilize oriented fuzzing to explore diverging paths from vulnerability point. For userspace programs, we adopt a control-flow stitching solution to stitch crashing paths and diverging paths together to generate exploit. For kernel UAF, we leverage a lightweight symbolic execution to identify, analyze and evaluate the system calls valuable and useful for exploiting vulnerabilities. We have developed a prototype system and evaluated it on a set of 19 CTF (capture the flag) programs and 15 realworld Linux kernel UAF vulnerabilities. Experiment results showed it could generate exploit for most of the userspace test set, and it could also facilitate security mitigation bypassing and exploitability evaluation for kernel test set. Wei Wu 0010, Chao Zhang 0008, Xinyu Xing 0001, Xiaorui Gong |
Cybersecur. | 4 |
| 2019 | Building a Trustworthy Execution Environment to Defeat Exploits from both Cyber Space and Physical Space for ARMabstractThe rapid evolution of Internet-of-Things (IoT) technologies has led to an emerging need to make them smarter. However, the smartness comes at the cost of multi-vector security exploits. From cyber space, a compromised operating system could access all the data in a cloud-aware IoT device. From physical space, cold-boot attacks and DMA attacks impose a great threat to the unattended devices. In this paper, we propose TrustShadow that provides a comprehensively protected execution environment for unmodified application running on ARM-based IoT devices. To defeat cyber attacks, TrustShadow takes advantage of ARM TrustZone technology and partitions resources into the secure and normal worlds. In the secure world, TrustShadow constructs a trusted execution environment for security-critical applications. This trusted environment is maintained by a lightweight runtime system. The runtime system does not provide system services itself. Rather, it forwards them to the untrusted normal-world OS, and verifies the returns. The runtime system further employs a page based encryption mechanism to ensure that all the data segments of a security-critical application appear in ciphertext in DRAM chip. When an encrypted data page is accessed, it is transparently decrypted to a page in the internal RAM, which is immune to physical exploits. Le Guan, Chen Cao 0004, Peng Liu 0005, Xinyu Xing 0001, Xinyang Ge, Shengzhi Zhang, Meng Yu 0001, Trent Jaeger |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2018 | LEMNA: Explaining Deep Learning based Security ApplicationsabstractWhile deep learning has shown a great potential in various domains, the lack of transparency has limited its application in security or safety-critical areas. Existing research has attempted to develop explanation techniques to provide interpretable explanations for each classification decision. Unfortunately, current methods are optimized for non-security tasks ( e.g., image analysis). Their key assumptions are often violated in security applications, leading to a poor explanation fidelity. In this paper, we propose LEMNA, a high-fidelity explanation method dedicated for security applications. Given an input data sample, LEMNA generates a small set of interpretable features to explain how the input sample is classified. The core idea is to approximate a local area of the complex deep learning decision boundary using a simple interpretable model. The local interpretable model is specially designed to (1) handle feature dependency to better work with security applications ( e.g., binary code analysis); and (2) handle nonlinear local boundaries to boost explanation fidelity. We evaluate our system using two popular deep learning applications in security (a malware classifier, and a function start detector for binary reverse-engineering). Extensive evaluations show that LEMNA's explanation has a much higher fidelity level compared to existing methods. In addition, we demonstrate practical use cases of LEMNA to help machine learning developers to validate model behavior, troubleshoot classification errors, and automatically patch the errors of the target models. Wenbo Guo 0002, Dongliang Mu, Jun Xu 0024, Purui Su, Gang Wang 0011, Xinyu Xing 0001 |
CCS | 6 |
| 2018 | Defending Against Adversarial Samples Without Security through ObscurityabstractIt has been recently shown that deep neural networks (DNNs) are susceptible to a particular type of attack that exploits a fundamental flaw in their design. This attack consists of generating particular synthetic examples referred to as adversarial samples. These samples are constructed by slightly manipulating real data-points that change "fool" the original DNN model, forcing it to misclassify previously correctly classified samples with high confidence. Many believe addressing this flaw is essential for DNNs to be used in critical applications such as cyber security. Previous work has shown that learning algorithms that enhance the robustness of DNN models all use the tactic of "security through obscurity". This means that security can be guaranteed only if one can obscure the learning algorithms from adversaries. Once the learning technique is disclosed, DNNs protected by these defense mechanisms are still susceptible to adversarial samples. In this work, we investigate by examining how previous research dealt with this and propose a generic approach to enhance a DNN's resistance to adversarial samples. More specifically, our approach integrates a data transformation module with a DNN, making it robust even if we reveal the underlying learning algorithm. To demonstrate the generality of our proposed approach and its potential for handling cyber security applications, we evaluate our method and several other existing solutions on datasets publicly available, such as a large scale malware dataset and MNIST and IMDB datasets. Our results indicate that our approach typically provides superior classification performance and robustness to attacks compared with state-of-art solutions. Wenbo Guo 0002, Qinglong Wang 0003, Kaixuan Zhang 0002, Alexander Ororbia, Sui Huang, Xue (Steve) Liu, C. Lee Giles, Lin Lin 0003, Xinyu Xing 0001 |
ICDM | 9 |
| 2018 | Explaining Deep Learning Models - A Bayesian Non-parametric ApproachabstractUnderstanding and interpreting how machine learning (ML) models make decisions have been a big challenge. While recent research has proposed various technical approaches to provide some clues as to how an ML model makes individual predictions, they cannot provide users with an ability to inspect a model as a complete entity. In this work, we propose a novel technical approach that augments a Bayesian non-parametric regression mixture model with multiple elastic nets. Using the enhanced mixture model, we can extract generalizable insights for a target model through a global approximation. To demonstrate the utility of our approach, we evaluate it on different ML models in the context of image recognition. The empirical results indicate that our proposed approach not only outperforms the state-of-the-art techniques in explaining individual decisions but also provides users with an ability to discover the vulnerabilities of the target ML models. Wenbo Guo 0002, Sui Huang, Yunzhe Tao, Xinyu Xing 0001, Lin Lin 0003 |
NeurIPS | 4 |
| 2018 | Understanding the Reproducibility of Crowd-reported Security Vulnerabilities
Dongliang Mu, Alejandro Cuevas, Hang Hu 0002, Xinyu Xing 0001, Bing Mao 0001, Gang Wang 0011 |
USENIX Security Symposium | 5 |
| 2018 | FUZE: Towards Facilitating Exploit Generation for Kernel Use-After-Free Vulnerabilities
Wei Wu 0010, Yueqi Chen 0001, Jun Xu 0024, Xinyu Xing 0001, Xiaorui Gong |
USENIX Security Symposium | 4 |
| 2018 | An Empirical Evaluation of Rule Extraction from Recurrent Neural NetworksabstractRule extraction from black box models is critical in domains that require model validation before implementation, as can be the case in credit scoring and medical diagnosis. Though already a challenging problem in statistical learning in general, the difficulty is even greater when highly nonlinear, recursive models, such as recurrent neural networks (RNNs), are fit to data. Here, we study the extraction of rules from second-order RNNs trained to recognize the Tomita grammars. We show that production rules can be stably extracted from trained RNNs and that in certain cases, the rules outperform the trained RNNs. Qinglong Wang 0003, Kaixuan Zhang 0002, Alexander Ororbia, Xinyu Xing 0001, Xue (Steve) Liu, C. Lee Giles |
Neural Comput. | 4 |
| 2017 | Supporting Transparent Snapshot for Bare-metal Malware Analysis on Mobile DevicesabstractThe increasing growth of cybercrimes targeting mobile devices urges an efficient malware analysis platform. With the emergence of evasive malware, which is capable of detecting that it is being analyzed in virtualized environments, bare-metal analysis has become the definitive resort. Existing works mainly focus on extracting the malicious behaviors exposed during bare-metal analysis. However, after malware analysis, it is equally important to quickly restore the system to a clean state to examine the next sample. Unfortunately, state-of-the-art solutions on mobile platforms can only restore the disk, and require a time-consuming system reboot. In addition, all of the existing works require some in-guest components to assist the restoration. Therefore, a kernel-level malware is still able to detect the presence of the in-guest components. Le Guan, Shijie Jia 0001, Bo Chen 0028, Fengwei Zhang, Bo Luo, Jingqiang Lin 0001, Peng Liu 0005, Xinyu Xing 0001, Luning Xia |
ACSAC | 8 |
| 2017 | FlashGuard: Leveraging Intrinsic Flash Properties to Defend Against Encryption RansomwareabstractEncryption ransomware is a malicious software that stealthily encrypts user files and demands a ransom to provide access to these files. Several prior studies have developed systems to detect ransomware by monitoring the activities that typically occur during a ransomware attack. Unfortunately, by the time the ransomware is detected, some files already undergo encryption and the user is still required to pay a ransom to access those files. Furthermore, ransomware variants can obtain kernel privilege, which allows them to terminate software-based defense systems, such as anti-virus. While periodic backups have been explored as a means to mitigate ransomware, such backups incur storage overheads and are still vulnerable as ransomware can obtain kernel privilege to stop or destroy backups. Ideally, we would like to defend against ransomware without relying on software-based solutions and without incurring the storage overheads of backups. Jian Huang 0006, Jun Xu 0024, Xinyu Xing 0001, Peng Liu 0005, Moinuddin K. Qureshi |
CCS | 3 |
| 2017 | What You See is Not What You Get! Thwarting Just-in-Time ROP with ChameleonabstractAddress space randomization has long been used for counteracting code reuse attacks, ranging from conventional ROP to sophisticated Just-in-Time ROP. At the high level, it shuffles program code in memory and thus prevents malicious ROP payload from performing arbitrary operations. While effective in mitigating attacks, existing randomization mechanisms are impractical for real-world applications and systems, especially considering the significant performance overhead and potential program corruption incurred by their implementation. In this paper, we introduce CHAMELEON, a practical defense mechanism that hinders code reuse attacks, particularly Just-in-Time ROP attacks. Technically speaking, CHAMELEON instruments program code, randomly shuffles code page addresses and minimizes the attack surface exposed to adversaries. While this defense mechanism follows in the footprints of address space randomization, our design principle focuses on using randomization to obstruct code page disclosure, making the ensuing attacks infeasible. We implemented a prototype of CHAMELEON on Linux operating system and extensively experimented it in different settings. Our theoretical and empirical evaluation indicates the effectiveness and efficiency of CHAMELEON in thwarting Just-in-Time ROP attacks. Ping Chen 0003, Jun Xu 0024, Zhisheng Hu, Xinyu Xing 0001, Bing Mao 0001, Peng Liu 0005 |
DSN | 4 |
| 2017 | Adversary Resistant Deep Neural Networks with an Application to Malware DetectionabstractOutside the highly publicized victories in the game of Go, there have been numerous successful applications of deep learning in the fields of information retrieval, computer vision, and speech recognition. In cybersecurity, an increasing number of companies have begun exploring the use of deep learning (DL) in a variety of security tasks with malware detection among the more popular. These companies claim that deep neural networks (DNNs) could help turn the tide in the war against malware infection. However, DNNs are vulnerable to adversarial samples, a shortcoming that plagues most, if not all, statistical and machine learning models. Recent research has demonstrated that those with malicious intent can easily circumvent deep learning-powered malware detection by exploiting this weakness. Qinglong Wang 0003, Wenbo Guo 0002, Kaixuan Zhang 0002, Alexander Ororbia, Xinyu Xing 0001, Xue (Steve) Liu, C. Lee Giles |
KDD | 5 |
| 2017 | TrustShadow: Secure Execution of Unmodified Applications with ARM TrustZoneabstractThe rapid evolution of Internet-of-Things (IoT) technologies has led to an emerging need to make them smarter. A variety of applications now run simultaneously on an ARM-based processor. For example, devices on the edge of the Internet are provided with higher horsepower to be entrusted with storing, processing and analyzing data collected from IoT devices. This significantly improves efficiency and reduces the amount of data that needs to be transported to the cloud for data processing, analysis and storage. However, commodity OSes are prone to compromise. Once they are exploited, attackers can access the data on these devices. Since the data stored and processed on the devices can be sensitive, left untackled, this is particularly disconcerting. In this paper, we propose a new system, TrustShadow that shields legacy applications from untrusted OSes. TrustShadow takes advantage of ARM TrustZone technology and partitions resources into the secure and normal worlds. In the secure world, TrustShadow constructs a trusted execution environment for security-critical applications. This trusted environment is maintained by a lightweight runtime system that coordinates the communication between applications and the ordinary OS running in the normal world. The runtime system does not provide system services itself. Rather, it forwards requests for system services to the ordinary OS, and verifies the correctness of the responses. To demonstrate the efficiency of this design, we prototyped TrustShadow on a real chip board with ARM TrustZone support, and evaluated its performance using both microbenchmarks and real-world applications. We showed TrustShadow introduces only negligible overhead to real-world applications. Le Guan, Peng Liu 0005, Xinyu Xing 0001, Xinyang Ge, Shengzhi Zhang, Meng Yu 0001, Trent Jaeger |
MobiSys | 3 |
| 2017 | System Service Call-oriented Symbolic Execution of Android Framework with Applications to Vulnerability Discovery and Exploit GenerationabstractAndroid Application Framework is an integral and foundational part of the Android system. Each of the 1.4 billion Android devices relies on the system services of Android Framework to manage applications and system resources. Given its critical role, a vulnerability in the framework can be exploited to launch large-scale cyber attacks and cause severe harms to user security and privacy. Recently, many vulnerabilities in Android Framework were exposed, showing that it is vulnerable and exploitable. However, most of the existing research has been limited to analyzing Android applications, while there are very few techniques and tools developed for analyzing Android Framework. In particular, to our knowledge, there is no previous work that analyzes the framework through symbolic execution, an approach that has proven to be very powerful for vulnerability discovery and exploit generation. We design and build the first system, Centaur, that enables symbolic execution of Android Framework. Due to some unique characteristics of the framework, such as its middleware nature and extraordinary complexity, many new challenges arise and are tackled in Centaur. In addition, we demonstrate how the system can be applied to discovering new vulnerability instances, which can be exploited by several recently uncovered attacks against the framework, and to generating PoC exploits. Lannan Luo, Qiang Zeng 0001, Chen Cao 0004, Kai Chen 0012, Jian Liu 0008, Neng Gao, Min Yang 0002, Xinyu Xing 0001, Peng Liu 0005 |
MobiSys | 9 |
| 2017 | Postmortem Program Analysis with Hardware-Enhanced Post-Crash Artifacts
Jun Xu 0024, Dongliang Mu, Xinyu Xing 0001, Peng Liu 0005, Ping Chen 0003, Bing Mao 0001 |
USENIX Security Symposium | 3 |
| 2016 | CREDAL: Towards Locating a Memory Corruption Vulnerability with Your Core DumpabstractAfter a program has crashed and terminated abnormally, it typically leaves behind a snapshot of its crashing state in the form of a core dump. While a core dump carries a large amount of information, which has long been used for software debugging, it barely serves as informative debugging aids in locating software faults, particularly memory corruption vulnerabilities. A memory corruption vulnerability is a special type of software faults that an attacker can exploit to manipulate the content at a certain memory. As such, a core dump may contain a certain amount of corrupted data, which increases the difficulty in identifying useful debugging information (e.g. , a crash point and stack traces). Without a proper mechanism to deal with this problem, a core dump can be practically useless for software failure diagnosis. In this work, we develop CREDAL, an automatic tool that employs the source code of a crashing program to enhance core dump analysis and turns a core dump to an informative aid in tracking down memory corruption vulnerabilities. Specifically, CREDAL systematically analyzes a core dump potentially corrupted and identifies the crash point and stack frames. For a core dump carrying corrupted data, it goes beyond the crash point and stack trace. In particular, CREDAL further pinpoints the variables holding corrupted data using the source code of the crashing program along with the stack frames. To assist software developers (or security analysts) in tracking down a memory corruption vulnerability, CREDAL also performs analysis and highlights the code fragments corresponding to data corruption. Jun Xu 0024, Dongliang Mu, Ping Chen 0003, Xinyu Xing 0001, Pei Wang 0007, Peng Liu 0005 |
CCS | 4 |
| 2016 | From Physical to Cyber: Escalating Protection for Personalized Auto InsuranceabstractNowadays, auto insurance companies set personalized insurance rate based on data gathered directly from their customers' cars. In this paper, we show such a personalized insurance mechanism -- wildly adopted by many auto insurance companies -- is vulnerable to exploit. In particular, we demonstrate that an adversary can leverage off-the-shelf hardware to manipulate the data to the device that collects drivers' habits for insurance rate customization and obtain a fraudulent insurance discount. In response to this type of attack, we also propose a defense mechanism that escalates the protection for insurers' data collection. The main idea of this mechanism is to augment the insurer's data collection device with the ability to gather unforgeable data acquired from the physical world, and then leverage these data to identify manipulated data points. Our defense mechanism leveraged a statistical model built on unmanipulated data and is robust to manipulation methods that are not foreseen previously. We have implemented this defense mechanism as a proof-of-concept prototype and tested its effectiveness in the real world. Our evaluation shows that our defense mechanism exhibits a false positive rate of 0.032 and a false negative rate of 0.013. Le Guan, Jun Xu 0024, Shuai Wang 0011, Xinyu Xing 0001, Lin Lin 0003, Heqing Huang 0001, Peng Liu 0005, Wenke Lee |
SenSys | 4 |
| 2016 | WebRanz: web page randomization for better advertisement delivery and web-bot preventionabstractNowadays, a rapidly increasing number of web users are using Ad-blockers to block online advertisements. Ad-blockers are browser-based software that can block most Ads on the websites, speeding up web browsers and saving bandwidth. Despite these benefits to end users, Ad-blockers could be catastrophic for the economic structure underlying the web, especially considering the rise of Ad-blocking as well as the number of technologies and services that rely exclusively on Ads to compensate their cost. In this paper, we introduce WebRanz that utilizes a randomization mechanism to circumvent Ad-blocking. Using WebRanz, content publishers can constantly mutate the internal HTML elements and element attributes of their web pages, without affecting their visual appearances and functionalities. Randomization invalidates the pre-defined patterns that Ad-blockers use to filter out Ads. Though the design of WebRanz is motivated by evading Ad-blockers, WebRanz also benefits the defense against bot scripts. We evaluate the effectiveness of WebRanz and its overhead using 221 randomly sampled top Alexa web pages and 8 representative bot scripts. Weihang Wang 0001, Yunhui Zheng, Xinyu Xing 0001, Yonghwi Kwon 0001, Xiangyu Zhang 0001, Patrick Eugster |
SIGSOFT FSE | 3 |
| 2016 | TrackMeOrNot: Enabling Flexible Control on Web TrackingabstractRecent advance in web tracking technologies has raised many privacy concerns. To combat users' fear of privacy invasion, online vendors have taken measures such as being more transparent with users about their data use and providing options for users to manage their online activities. Such efforts gain users' trust in online vendors and improve their willingness to share their digital footprints. However, there are still a significant amount of users who actively limit involuntarily sharing of data because vendor provided management tools only restrict the use of collected data and users worry vendors do not have enough measures in place to protect their privacy sensitive information. Wei Meng 0001, Byoungyoung Lee, Xinyu Xing 0001, Wenke Lee |
WWW | 3 |
| 2015 | UCognito: Private Browsing without TearsabstractWhile private browsing is a standard feature, its implementation has been inconsistent among the major browsers. More seriously, it often fails to provide the adequate or even the intended privacy protection. For example, as shown in prior research, browser extensions and add-ons often undermine the goals of private browsing. In this paper, we first present our systematic study of private browsing. We developed a technical approach to identify browser traces left behind by a private browsing session, and showed that Chrome and Firefox do not correctly clear some of these traces. We analyzed the source code of these browsers and discovered that the current implementation approach is to decide the behaviors of a browser based on the current browsing mode (i.e., private or public); but such decision points are scattered throughout the code base. This implementation approach is very problematic because developers are prone to make mistakes given the complexities of browser components (including extensions and add-ons). Based on this observation, we propose a new and general approach to implement private browsing. The main idea is to overlay the actual filesystem with a sandbox filesystem when the browser is in private browsing mode, so that no unintended leakage is allowed and no persistent modification is stored. This approach requires no change to browsers and the OS kernel because the layered sandbox filesystem is implemented by interposing system calls. We have implemented a prototype system called Ucognito on Linux. Our evaluations show that Ucognito, when applied to Chrome and Firefox, stops all known privacy leaks identified by prior work and our current study. More importantly, Ucognito incurs only negligible performance overhead: e.g., 0%-2.5% in benchmarks for standard JavaScript and webpage loading. Meng Xu 0001, Yeongjin Jang, Xinyu Xing 0001, Taesoo Kim, Wenke Lee |
CCS | 3 |
| 2015 | Understanding Malvertising Through Ad-Injecting Browser ExtensionsabstractMalvertising is a malicious activity that leverages advertising to distribute various forms of malware. Because advertising is the key revenue generator for numerous Internet companies, large ad networks, such as Google, Yahoo and Microsoft, invest a lot of effort to mitigate malicious ads from their ad networks. This drives adversaries to look for alternative methods to deploy malvertising. In this paper, we show that browser extensions that use ads as their monetization strategy often facilitate the deployment of malvertising. Moreover, while some extensions simply serve ads from ad networks that support malvertising, other extensions maliciously alter the content of visited webpages to force users into installing malware. To measure the extent of these behaviors we developed Expector, a system that automatically inspects and identifies browser extensions that inject ads, and then classifies these ads as malicious or benign based on their landing pages. Using Expector, we automatically inspected over 18,000 Chrome browser extensions. We found 292 extensions that inject ads, and detected 56 extensions that participate in malvertising using 16 different ad networks and with a total user base of 602,417. Xinyu Xing 0001, Wei Meng 0001, Byoungyoung Lee, Udi Weinsberg, Anmol Sheth, Roberto Perdisci, Wenke Lee |
WWW | 1 |
| 2015 | ISC: An Iterative Social Based Classifier for Adult Account Detection on TwitterabstractThe widespread of adult content on online social networks (e.g., Twitter) is becoming an emerging yet critical problem. An automatic method to identify accounts spreading sexually explicit content (i.e., adult account) is of significant values in protecting children and improving user experiences. Traditional adult content detection techniques are ill-suited for detecting adult accounts on Twitter due to the diversity and dynamics in Twitter content. In this paper, we formulate the adult account detection as a graph based classification problem and demonstrate our detection method on Twitter by using social links between Twitter accounts and entities in tweets. As adult Twitter accounts are mostly connected with normal accounts and post many normal entities, which makes the graph full of noisy links, existing graph based classification techniques cannot work well on such a graph. To address this problem, we propose an iterative social based classifier (ISC), a novel graph based classification technique resistant to the noisy links. Evaluations using large-scale real-world Twitter data show that, by labeling a small number of popular Twitter accounts, ISC can achieve satisfactory performance in adult account detection, significantly outperforming existing techniques. Hanqiang Cheng, Xinyu Xing 0001, Xue (Steve) Liu, Qin Lv |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2014 | Your Online Interests: Pwned! A Pollution Attack Against Targeted AdvertisingabstractWe present a new ad fraud mechanism that enables publishers to increase their ad revenue by deceiving the ad exchange and advertisers to target higher paying ads at users visiting the publisher's site. Our attack is based on polluting users' online interest profile by issuing requests to content not explicitly requested by the user, such that it influences the ad selection process. We address several challenges involved in setting up the attack for the two most commonly used ad targeting mechanisms -- re-marketing and behavioral targeting. We validate the attack for one of the largest ad exchanges and empirically measure the monetary gains of the publisher by emulating the attack using web traces of 619 real users. Our results show that the attack is effective in biasing ads towards the desired higher-paying advertisers; the polluter can influence up to 74% and 12% of the total ad impressions for re-marketing and behavioral pollution, respectively. The attack is robust to diverse browsing patterns and online interests of users. Finally, the attack is lucrative and on average the attack can increase revenue of fraudlent publishers by as much as 33%. Wei Meng 0001, Xinyu Xing 0001, Anmol Sheth, Udi Weinsberg, Wenke Lee |
CCS | 2 |
| 2014 | Exposing Inconsistent Web Search Results with Bobble
Xinyu Xing 0001, Wei Meng 0001, Dan Doozan, Nick Feamster, Wenke Lee, Alex C. Snoeren |
PAM | 1 |
| 2013 | Take This Personally: Pollution Attacks on Personalized Services
Xinyu Xing 0001, Wei Meng 0001, Dan Doozan, Alex C. Snoeren, Nick Feamster, Wenke Lee |
USENIX Security Symposium | 1 |
| 2013 | SafeVchat: A System for Obscene Content Detection in Online Video Chat ServicesabstractOnline video chat services such as Chatroulette, Omegle, and vChatter that randomly match pairs of users in video chat sessions are quickly becoming very popular, with over a million users per month in the case of Chatroulette. A key problem encountered in such systems is the presence of flashers and obscene content. This problem is especially acute given the presence of underage minors in such systems. This article presents SafeVchat, a novel solution to the problem of flasher detection that employs an array of image detection algorithms. A key contribution of the article concerns how the results of the individual detectors are fused together into an overall decision classifying a user as misbehaving or not, based on Dempster-Shafer theory. The article introduces a novel, motion-based skin detection method that achieves significantly higher recall and better precision. The proposed methods have been evaluated over real-world data and image traces obtained from Chatroulette.com. SafeVchat has been deployed in Chatroulette. A combination of SafeVchat with human moderation has resulted in banning as many as 50,000 inappropriate users per day on Chatoulette. Furthermore, offensive content on Chatoulette has dropped significantly from 33.08% (before SafeVchat installation) to 3.49% (after SafeVchat installation). Yu-Li Liang, Xinyu Xing 0001, Hanqiang Cheng, Jianxun Dang, Sui Huang, Richard Han 0001, Xue (Steve) Liu, Qin Lv, Shivakant Mishra |
ACM Trans. Internet Techn. | 2 |
| 2012 | Scalable misbehavior detection in online video chat servicesabstractThe need for highly scalable and accurate detection and filtering of misbehaving users and obscene content in online video chat services has grown as the popularity of these services has exploded in popularity. This is a challenging problem because processing large amounts of video is compute intensive, decisions about whether a user is misbehaving or not must be made online and quickly, and moreover these video chats are characterized by low quality video, poorly lit scenes, diversity of users and their behaviors, diversity of the content, and typically short sessions. This paper presents EMeralD, a highly scalable system for accurately detecting and filtering misbehaving users in online video chat applications. EMeralD substantially improves upon the state-of-the-art filtering mechanisms by achieving much lower computational cost and higher accuracy. We demonstrate EMeralD's improvement via experimental evaluations on real-world data sets obtained from Chatroulette.com. Xinyu Xing 0001, Yu-Li Liang, Sui Huang, Hanqiang Cheng, Richard Han 0001, Qin Lv, Xue (Steve) Liu, Shivakant Mishra, Yi Zhu 0010 |
KDD | 1 |
| 2012 | Demo: MVChat: flasher detection for mobile video chatabstractOnline video chat services such as Chatroulette [1] and Omegle [2] that randomly match pairs of users in video chat sessions have become increasingly popular, with over twenty thousand online users at anytime during a day. A key problem encountered in such systems is the presence of misbehaving users ("flashers") and obscene content. Our previous works [3] [4] prove that using some image recognition methods (skin-detection, dense SIFT) and machine learning algorithms could achieve significantly higher recall and better precision for flasher detection. Nowadays, with the rapid development of advanced mobile phones with both front and back cameras, we expect mobile video chat to become a popular extension of online video chat services. However, because of the computation-intensive features used by our previous solutions and mobile phones' hardware limitations such as memory size and CPU capacity, it is difficult to directly apply our previous works to mobile platforms. As smartphones are increasingly equipped with diverse sensing capabilities, we plan to utilize this multi-dimensional sensor information to extend flasher detection on mobile platform. This project explores how we can mine accelerometer and other mobile sensor data to infer some clues to optimize flasher detection accuracy while reducing the computation demands of flasher detection on the mobile device. Lei Tian 0004, Junho Ahn, Hanqiang Cheng, Xinyu Xing 0001, Yu-Li Liang, Shivakant Mishra, David Chu, Xue (Steve) Liu, Richard Han 0001, Qin Lv |
MobiSys | 4 |
| 2012 | Efficient misbehaving user detection in online video chat servicesabstractOnline video chat services, such as Chatroulette, Omegle, and vChatter are becoming increasingly popular and have attracted millions of users. One critical problem encountered in such applications is the presence of misbehaving users ("flashers") and obscene content. Automatically filtering out obscene content from these systems in an efficient manner poses a difficult challenge. This paper presents a novel Fine-Grained Cascaded (FGC) classification solution that significantly speeds up the compute-intensive process of classifying misbehaving users by dividing image feature extraction into multiple stages and filtering out easily classified images in earlier stages, thus saving unnecessary computation costs of feature extraction in later stages. Our work is further enhanced by integrating new webcam-related contextual information (illumination and color) into the classification process, and a 2-stage soft margin SVM algorithm for combining multiple features. Evaluation results using real-world data set obtained from Chatroulette show that the proposed FGC based classification solution significantly outperforms state-of-the-art techniques. Hanqiang Cheng, Yu-Li Liang, Xinyu Xing 0001, Xue (Steve) Liu, Richard Han 0001, Qin Lv, Shivakant Mishra |
WSDM | 3 |
| 2011 | A highly scalable bandwidth estimation of commercial hotspot access pointsabstractWiFi access points that provide Internet access to users have been steadily increasing in urban areas. Different access points differ from one another in terms of services that they provide, including available upstream and downstream bandwidths, overall network capacity, open/blocked ports, security features, and so on. However, there is no reliable service available at present that can aid a user in selecting an access point from the many that are available. The primary research challenge is how to accurately estimate the current backhaul bandwidth of different access points in an efficient manner without requiring any installation of special software on the access points, and not burdening the WiFi subscribers to perform any communication or computation intensive task. This paper presents a new highly scalable bandwidth estimation technique that is suitable for efficiently estimating the backhaul bandwidth of a large number of APs. This technique has been extensively evaluated via a prototype implementation in an indoor testbed and in the Amazon EC2 platform. The evaluation demonstrates that the proposed technique exhibits high measurement accuracy, low latency, high scalability, and minimal intrusiveness. Xinyu Xing 0001, Jianxun Dang, Shivakant Mishra, Xue (Steve) Liu |
INFOCOM | 1 |
| 2011 | SafeVchat: detecting obscene content and misbehaving users in online video chat servicesabstractOnline video chat services such as Chatroulette, Omegle, and vChatter that randomly match pairs of users in video chat sessions are fast becoming very popular, with over a million users per month in the case of Chatroulette. A key problem encountered in such systems is the presence of flashers and obscene content. This problem is especially acute given the presence of underage minors in such systems. This paper presents SafeVchat, a novel solution to the problem of flasher detection that employs an array of image detection algorithms. A key contribution of the paper concerns how the results of the individual detectors are fused together into an overall decision classifying the user as misbehaving or not, based on Dempster-Shafer Theory. The paper introduces a novel, motion-based skin detection method that achieves significantly higher recall and better precision. The proposed methods have been evaluated over real-world data and image traces obtained from Chatroulette.com. Xinyu Xing 0001, Yu-Li Liang, Hanqiang Cheng, Jianxun Dang, Sui Huang, Richard Han 0001, Xue (Steve) Liu, Qin Lv, Shivakant Mishra |
WWW | 1 |
| 2010 | Enhancing group recommendation by incorporating social relationship interactionsabstractGroup recommendation, which makes recommendations to a group of users instead of individuals, has become increasingly important in both the workspace and people’s social activities, such as brainstorming sessions for coworkers and social TV for family members or friends. Group recommendation is a challenging problem due to the dynamics of group memberships and diversity of group members. Previous work focused mainly on the content interests of group members and ignored the social characteristics within a group, resulting in suboptimal group recommendation performance. In this work, we propose a group recommendation method that utilizes both social and content interests of group members. We study the key characteristics of groups and propose (1) a group consensus function that captures the social, expertise, and interest dissimilarity among multiple group members; and (2) a generic framework that automatically analyzes group characteristics and constructs the corresponding group consensus function. Detailed user studies of diverse groups demonstrate the effectiveness of the proposed techniques, and the importance of incorporating both social and content interests in group recommender systems. Mike Gartrell, Xinyu Xing 0001, Qin Lv, Aaron Beach, Richard Han 0001, Shivakant Mishra, Karim Seada |
GROUP | 2 |
| 2010 | ARBOR: Hang Together Rather Than Hang Separately in 802.11 WiFi NetworksabstractWith 802.11 WiFi networks becoming popular in homes, it is common for an end-user to have access to multiple WiFi access points (APs) from residents next door. In general, wireless networks have much higher bandwidth than residential Internet (DSL or Cable) connections. This provides an incentive for an end-user to simultaneously harness bandwidths from multiple APs. This paper introduces ARBOR, an 802.11 driver that aggregates broadband connections in a neighborhood and maximizes Internet access bandwidth in a secure manner. ARBOR has four important characteristics. First, ARBOR can sustain a much longer switching cycle without losing packets queued at different APs. Second, it can schedule traffic loads and (in)directly aggregate AP backhaul bandwidths. Third, ARBOR designs and implements a light-weight authentication mechanism that provides sufficient amount of security, and at the same time, ensures fast switching time. Finally, ARBOR is transparent to the upper layers of the network stack. A prototype of ARBOR has been implemented and extensively evaluated. Experiment results show that ARBOR provides significantly better throughput gains and lower Internet access delays. Xinyu Xing 0001, Shivakant Mishra, Xue (Steve) Liu |
INFOCOM | 1 |
| 2009 | Where is the tight link in a home wireless broadband environment?abstractNowadays, tariffs and packages on offer for broadband worldwide gives a general story of increased speed and reduced prices. As a complementary to fixed-line broadband access, WiFi networks are undoubtedly taking off in many American households. Due to the combination of wireless technology and fixed-line broadband access, prior available bandwidth measurement tools may not be an appropriate solution for a common broadband subscriber. In this paper, we utilize abget in combination with Multi-Router Traffic Grapher (MRTG) to target the tight link of the Internet. Based on the observation that the tight link of the Internet in the context of an in-home wireless broadband access is usually on the edge of the Internet, we then introduce ABODE, a single-end, light-weight tool for estimating available bandwidth in the context of an in-home wireless bandwidth environment. ABODE harnesses ICMP echo request messages to generate stable, rate-controlled traffic flows. Based on the timestamp of ICMP echo reply messages, ABODE performs available bandwidth estimation by calculating the drift of the time centroid over a traffic flow. To verify the accuracy of ABODE, we use MRTG data to compare with the estimation results of ABODE. The measurement shows that ABODE is capable of efficiently estimating available bandwidth in the context of an in-home wireless broadband environment. Xinyu Xing 0001, Shivakant Mishra |
MASCOTS | 1 |