VLDB 2026 Research / reviewers in the wild / expert
Yuan Tian 0001
dblp:39/5423-1
· DBLP profile ↗
58ranked-venue papers
3as first author
39since 2021 · last 2026
0000-0002-6435-564XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 37 · 3 first-author · 23 since 2021Artificial intelligence and machine learning · 13 · 11 since 2021Software engineering, systems software and programming languages · 5 · 3 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GeoGen: A Two-stage Coarse-to-Fine Framework for Fine-grained Synthetic Location-based Social Network Trajectory GenerationabstractLocation-Based Social Network (LBSN) check-in trajectory data are important for many practical applications like POI recommendation, advertising, and pandemic intervention. However, the high collection costs and ever-increasing privacy concerns prevent us from accessing large-scale LBSN trajectory data. The recent advances in synthetic data generation provide us with a new opportunity to achieve this, which utilizes generative AI to generate synthetic data that preserves the characteristics of real data while ensuring privacy protection. However, generating synthetic LBSN check-in trajectories remains challenging due to their spatially discrete, temporally irregular nature and the complex spatio-temporal patterns caused by sparse activities and uncertain human mobility. To address this challenge, we propose GeoGen, a two-stage coarse-to-fine framework for large-scale LBSN check-in trajectory generation. In the first stage, we reconstruct spatially continuous, temporally regular latent movement sequences from the original LBSN check-in trajectories and then design a Sparsity-aware Spatio-temporal Diffusion model (S^2TDiff) with an efficient denosing network to learn their underlying behavioral patterns. In the second stage, we design Coarse2FineNet, a Transformer-based Seq2Seq architecture equipped with a dynamic context fusion mechanism in the encoder and a multi-task hybrid-head decoder, which generates fine-grained LBSN trajectories based on coarse-grained latent movement sequences by modeling semantic relevance and behavioral uncertainty. Extensive experiments on four real-world datasets show that GeoGen excels state-of-the-art models for both fidelity and utility evaluation, e.g., it increases over 69% and 55% in distance and radius metrics on the FS-TKY dataset. Rongchao Xu, Kunlin Cai, Lin Jiang 0007, Zhiqing Hong, Yuan Tian 0001, Guang Wang 0001 |
AAAI | 5 |
| 2026 | FineSteer: A Unified Framework for Fine-Grained Inference-Time Steering in Large Language ModelsabstractLarge language models (LLMs) often exhibit undesirable behaviors, such as safety violations and hallucinations.Although inference-time steering offers a cost-effective way to adjust model behavior without updating its parameters, existing methods often fail to be simultaneously effective, utilitypreserving, and training-efficient due to their rigid, one-size-fits-all designs and limited adaptability.In this work, we present FineSteer, a novel steering framework that decomposes inference-time steering into two complementary stages-conditional steering and fine-grained vector synthesis-allowing finegrained control over when and how to steer internal representations.In the first stage, we introduce a Subspace-guided Conditional Steering (SCS) mechanism that preserves model utility by avoiding unnecessary steering.In the second stage, we propose a Mixture-of-Steering-Experts (MoSE) mechanism that captures the multimodal nature of desired steering behaviors and generates query-specific steering vectors for improved effectiveness.Through tailored designs in both SCS and MoSE, FineSteer maintains robust performance on general queries while adaptively optimizing steering vectors for targeted inputs in a training-efficient manner.Extensive experiments on safety and truthfulness benchmarks show that FineSteer outperforms the state-of-the-art methods in overall performance (e.g., A 7.6% improvement on TruthfulQA over Llama-3), achieving stronger steering performance with minimal utility loss.The code is available at https://github.com/YukinoAsuna/FineSteer. Zixuan Weng, Jinghuai Zhang, Kunlin Cai, Ying Li 0095, Peiran Wang, Yuan Tian 0001 |
ACL (1) | 6 |
| 2026 | From Perception to Protection: A Developer-Centered Study of Security and Privacy Threats in Extended Reality (XR)
Kunlin Cai, Jinghuai Zhang, Ying Li 0095, Tianshi Li 0001, Yuan Tian 0001 |
NDSS | 7 |
| 2026 | Breaking the Illusion: Automated Reasoning of GDPR Consent Violations
Ying Li 0095, Wenjun Qiu, Faysal Hossain Shezan, Kunlin Cai, Michelangelo van Dam, Lisa M. Austin, David Lie, Yuan Tian 0001 |
SP | 8 |
| 2026 | Location-Enhanced Information Flow for Home AutomationsabstractSmart-home automations enable users to customize smart devices to react automatically to people, the environment, and more. For example, an automation might adjust the lights when people are at home or enable a garage door to open by voice command. While automations offer convenience and accessibility, they can also inadvertently expose users to security and privacy risks, such as leaking sensitive data or allowing untrusted parties to control users' devices. Prior work has shown that information flow analysis is a promising technique for identifying these kinds of risks, hypothesizing that the analysis would be yet more effective if it could differentiate between devices located in different places in the home. We tested this hypothesis by developing a tool that extends prior information flow analysis approaches to account for device location. We conducted an interview study with 22 participants to build a dataset of home automations to establish a ground truth to evaluate the tool. We found that incorporating device location leads to an improved analysis that identifies more of the vulnerabilities users care about (F1 score 0.74) compared to prior work (F1 score 0.29). Our results demonstrate the feasibility of incorporating device location into an information flow analysis and, perhaps more importantly, suggest additional ways to prevent security and privacy risks beyond controlling potentially unsafe information flows. McKenna McCall, Ben Weinshel, Kunlin Cai, Ying Li 0095, Eric Zeng 0001, Devika Manohar, Lujo Bauer, Limin Jia 0001, Yuan Tian 0001 |
Proc. Priv. Enhancing Technol. | 9 |
| 2025 | PromFuzz: Leveraging LLM-Driven and Bug-Oriented Composite Analysis for Detecting Functional Bugs in Smart ContractsabstractSmart contracts are fundamental pillars of the blockchain, playing a crucial role in facilitating various business transactions. However, these smart contracts are vulnerable to exploitable bugs that can lead to substantial monetary losses. A recent study reveals that over 80% of these exploitable bugs, which are primarily functional bugs, can evade the detection of current tools. Automatically identifying functional bugs in smart contracts presents challenges from multiple perspectives. The primary issue is the significant gap between understanding the high-level logic of the business model and checking the low-level implementations in smart contracts. Furthermore, identifying deeply rooted functional bugs in smart contracts requires the automated generation of effective detection oracles based on various bug features.To address these challenges, we design and implement PromFuzz, an automated and scalable system to detect functional bugs in smart contracts. In PromFuzz, we first propose a novel Large Language Model (LLM)-driven analysis framework, which leverages a dual-agent prompt engineering strategy to pinpoint potentially vulnerable functions for further scrutiny. We then implement a dual-stage coupling approach, which focuses on generating invariant checkers that leverage logic information extracted from potentially vulnerable functions. Finally, we design a bug-oriented fuzzing engine, which maps the logical information from the high-level business model to the low-level smart contract implementations, and performs the bug-oriented fuzzing on targeted functions. We evaluate PromFuzz from 4 perspectives on 5 ground-truth datasets and compare it with multiple state-of-the-art methods. The results show that PromFuzz achieves 86.96% recall and 93.02% F1-score in detecting functional bugs, marking at least a 50% improvement in both metrics over state-of-the-art methods. Moreover, we perform an in-depth analysis on 10 real-world DeFi projects and detect 30 zero-day bugs. Our further case studies, the risky first deposit bug and the AMM price oracle manipulation bug on real-world DeFi projects, demonstrate the serious risks of the exploitable functional bugs in smart contracts. Up to now, 24 zero-day bugs have been assigned CVE IDs. Our discoveries have safeguarded assets totaling $18.2 billion from potential monetary losses. Xingshuang Lin, Qinge Xie, Yuan Tian 0001, Saman A. Zonouz, Na Ruan, Raheem A. Beyah, Shouling Ji |
ASE | 4 |
| 2025 | Automated Repair of OpenID Connect ProgramsabstractOpenID Connect has revolutionized online authentication based on single sign-on (SSO) by providing a secure and convenient method for accessing multiple services with a single set of credentials. Despite its widespread adoption, critical security bugs in OpenID Connect have resulted in significant financial losses and security breaches, highlighting the need for robust mitigation strategies. Automated program repair presents a promising solution for generating candidate patches for OpenID implementations. However, challenges such as domain-specific complexities and the necessity for precise fault localization and patch verification must be addressed. We propose AuthFix, a counterexample-guided repair engine leveraging LLMs for automated OpenID bug fixing. AuthFix integrates three key components: fault localization, patch synthesis, and patch verification. By employing a novel Petri-net-based model checker, AuthFix ensures the correctness of patches by effectively modeling interactions. Our evaluation on a dataset of OpenID bugs demonstrates that AuthFix successfully generated correct patches for 17 out of 23 bugs (74%), with a high proportion of patches semantically equivalent to developer-written fixes. Tamjid Al Rahat, Yanju Chen, Yu Feng 0001, Yuan Tian 0001 |
ASE | 4 |
| 2025 | Firmrca: Towards Post-Fuzzing Analysis on ARM Embedded Firmware with Efficient Event-Based Fault LocalizationabstractWhile fuzzing has demonstrated its effectiveness in exposing vulnerabilities within embedded firmware, the discovery of crashing test cases is only the first step in improving the security of these critical systems. The subsequent fault localization process, which aims to precisely identify the root causes of observed crashes, is a crucial yet time-consuming post-fuzzing work. Unfortunately, the automated root cause analysis on embedded firmware crashes remains an underexplored area, which is challenging from several perspectives: (1) the fuzzing campaign towards the embedded firmware lacks adequate debugging mechanisms, making it hard to automatically extract essential runtime information for analysis; (2) the inherent raw binary nature of embedded firmware often leads to over-tainted and noisy suspicious instructions, which provides limited guidance for analysts in manually investigating the root cause and remediating the underlying vulnerability. To address these challenges, we design and implement FirmRCA, a practical fault localization framework tailored specifically for embedded firmware. FirmRCA introduces an event-based footprint collection approach that leverages concrete memory accesses in the crash reproducing process to aid and significantly expedite reverse execution. Next, to solve the complicated memory alias problem, FirmRCA proposes a history-driven method by tracking data propagation through the execution trace, enabling precise identification of deep crash origins. Finally, FirmRCA proposes a novel strategy to highlight key instructions related to the root cause, providing practical guidance in the final investigation. To demonstrate the efficacy of FirmRCA, we evaluate it with both synthetic and real-world targets, including 41 crashing test cases across 17 firmware images. The results show that FIRMRCA can effectively (92.7% success rate) identify the root cause of crashing test cases within the top 10 instructions. Compared to state-of-the-art works, FIRMRCA demonstrates its superiority in 27.8% improvement in full execution trace analysis capability, polynomial-level acceleration in overall efficiency and 73.2% higher success rate within the top 10 instructions in effectiveness. Boyu Chang, Peiyu Liu 0003, Yuan Tian 0001, Raheem A. Beyah, Shouling Ji |
SP | 5 |
| 2025 | SoK: Towards Effective Automated Vulnerability Repair
Ying Li 0095, Faysal Hossain Shezan, Bomin Wei, Gang Wang 0011, Yuan Tian 0001 |
USENIX Security Symposium | 5 |
| 2025 | Waltzz: WebAssembly Runtime Fuzzing with Stack-Invariant Transformation
Jiacheng Xu 0006, Peiyu Liu 0003, Qinge Xie, Yuan Tian 0001, Jianhai Chen, Shouling Ji |
USENIX Security Symposium | 6 |
| 2025 | A Magnetic Signal Based Device Fingerprinting Scheme in Wireless ChargingabstractWireless charging is widely used to charge smart devices with limited battery capacity. However, it is susceptible to the identity spoofing attack, where adversaries can impersonate malicious devices as legitimate ones to gain unauthorized access and potentially disrupt the wireless charging system (e.g., resulting in incorrect billing, overheating, or even explosions). Device fingerprinting is a classical method for defending against identity spoofing attacks. However, applying existing schemes in wireless charging scenarios has drawbacks such as inconvenience (e.g., requiring specialized devices or user participation) and ineffectiveness (e.g., vulnerability to spoofing). Thus, we design a novel passive, effective, and robust device fingerprinting scheme called MagID for wireless charging systems. The insight of MagID lies in the fact that during wireless charging, the magnetic signal around a device can reflect inherent hardware differences. These differences can be extracted as unique fingerprints for authentication purposes. MagID leverages a novel scheme, SUPER-ARRAY, to precisely measure magnetic data and generate effective fingerprints for authenticating a device's identity before starting charging progress. Experimental results demonstrate that MagID achieves an accuracy rate of 98.14% across various charging devices. We have also tested its performance under different impact factors and verified its compatibility with various wireless charging pads. Jiachun Li 0001, Yan Meng 0001, Guoxing Chen, Yuan Tian 0001, Haojin Zhu |
IEEE Trans. Dependable Secur. Comput. | 5 |
| 2024 | AuthSaber: Automated Safety Verification of OpenID Connect ProgramsabstractSingle Sign-On (SSO)-based authentication protocols, like OpenID Connect (OIDC), play a crucial role in enhancing security and privacy in today's interconnected digital world, gaining widespread adoption among the majority of prominent authentication service providers. These protocols establish a structured framework for verifying and authenticating the identities of individuals, organizations, and devices, while avoiding the necessity of sharing sensitive credentials (e.g., passwords) with external entities. However, the security guarantees of these protocols rely on their proper implementation, and real-world implementations can, and indeed often do, contain logical programming errors leading to severe attacks, including authentication bypass and user account takeover. In response to this challenge, we present AuthSaber, an automated verifier designed to assess the real-world OIDC protocol implementations against their standard safety specifications in a scalable manner. AuthSaber addresses the challenges of expressiveness for OIDC properties, modeling multi-party interactions, and automation by first designing a novel specification language based on linear temporal logic, leveraging an automaton-based approach to constrain the space of possible interactions between OIDC entities, and incorporating several domain-specific transformations to obtain programs and properties that can be directly reasoned about by software model checkers. We evaluate AuthSaber on the 15 most popular and widely used OIDC libraries and discover 16 previously unknown vulnerabilities, all of which are responsively disclosed to the developers. Five categories of these vulnerabilities also led to new CVEs. Tamjid Al Rahat, Yu Feng 0001, Yuan Tian 0001 |
CCS | 3 |
| 2024 | BadMerging: Backdoor Attacks Against Model MergingabstractFine-tuning pre-trained models for downstream tasks has led to a proliferation of open-sourced task-specific models. Recently, Model Merging (MM) has emerged as an effective approach to facilitate knowledge transfer among these independently fine-tuned models. MM directly combines multiple fine-tuned task-specific models into a merged model without additional training, and the resulting model shows enhanced capabilities in multiple tasks. Although MM provides great utility, it may come with security risks because an adversary can exploit MM to affect multiple downstream tasks. However, the security risks of MM have barely been studied. In this paper, we first find that MM, as a new learning paradigm, introduces unique challenges for existing backdoor attacks due to the merging process. To address these challenges, we introduce BadMerging, the first backdoor attack specifically designed for MM. Notably, BadMerging allows an adversary to compromise the entire merged model by contributing as few as one backdoored task-specific model. BadMerging comprises a two-stage attack mechanism and a novel feature-interpolation-based loss to enhance the robustness of embedded backdoors against the changes of different merging parameters. Considering that a merged model may incorporate tasks from different domains, BadMerging can jointly compromise the tasks provided by the adversary (on-task attack) and other contributors (off-task attack) and solve the corresponding unique challenges with novel attack designs. Extensive experiments show that BadMerging achieves remarkable attacks against various MM algorithms. Our ablation study demonstrates that the proposed attack designs can progressively contribute to the attack performance. Finally, we show that prior defense mechanisms fail to defend against our attacks, highlighting the need for more advanced defense. Our code is available at: https://github.com/jzhang538/BadMerging. Jinghuai Zhang, Jianfeng Chi, Zheng Li 0023, Kunlin Cai, Yang Zhang 0016, Yuan Tian 0001 |
CCS | 6 |
| 2024 | Where Have You Been? A Study of Privacy Risk for Point-of-Interest RecommendationabstractAs location-based services (LBS) have grown in popularity, more human mobility data has been collected. The collected data can be used to build machine learning (ML) models for LBS to enhance their performance and improve overall experience for users. However, the convenience comes with the risk of privacy leakage since this type of data might contain sensitive information related to user identities, such as home/work locations. Prior work focuses on protecting mobility data privacy during transmission or prior to release, lacking the privacy risk evaluation of mobility data-based ML models. To better understand and quantify the privacy leakage in mobility data-based ML models, we design a privacy attack suite containing data extraction and membership inference attacks tailored for point-of-interest (POI) recommendation models, one of the most widely used mobility data-based ML models. These attacks in our attack suite assume different adversary knowledge and aim to extract different types of sensitive information from mobility data, providing a holistic privacy risk assessment for POI recommendation models. Our experimental evaluation using two real-world mobility datasets demonstrates that current POI recommendation models are vulnerable to our attacks. We also present unique findings to understand what types of mobility data are more susceptible to privacy attacks. Finally, we evaluate defenses against these attacks and highlight future directions and challenges. Kunlin Cai, Jinghuai Zhang, Zhiqing Hong, William Shand, Guang Wang 0001, Desheng Zhang 0002, Jianfeng Chi, Yuan Tian 0001 |
KDD | 8 |
| 2024 | MOCK: Optimizing Kernel Fuzzing Mutation with Context-aware Dependency
Jiacheng Xu 0006, Xuhong Zhang 0002, Shouling Ji, Yuan Tian 0001, Qinying Wang, Peng Cheng 0001, Jiming Chen 0001 |
NDSS | 4 |
| 2024 | SyzTrust: State-aware Fuzzing on Trusted OS Designed for IoT DevicesabstractTrusted Execution Environments (TEEs) embedded in IoT devices provide a deployable solution to secure IoT applications at the hardware level. By design, in TEEs, the Trusted Operating System (Trusted OS) is the primary component. It enables the TEE to use security-based design techniques, such as data encryption and identity authentication. Once a Trusted OS has been exploited, the TEE can no longer ensure security. However, Trusted OSes for IoT devices have received little security analysis, which is challenging from several perspectives: (1) Trusted OSes are closed-source and have an unfavorable environment for sending test cases and collecting feedback. (2) Trusted OSes have complex data structures and require a stateful workflow, which limits existing vulnerability detection tools.To address the challenges, we present SyzTrust, the first state-aware fuzzing framework for vetting the security of resource-limited Trusted OSes. SyzTrust adopts a hardware-assisted framework to enable fuzzing Trusted OSes directly on IoT devices as well as tracking state and code coverage non-invasively. SyzTrust utilizes composite feedback to guide the fuzzer to effectively explore more states as well as to increase the code coverage. We evaluate SyzTrust on Trusted OSes from three major vendors: Samsung, Tsinglink Cloud, and Ali Cloud. These systems run on Cortex M23/33 MCUs, which provide the necessary abstraction for embedded TEEs. We discovered 70 previously unknown vulnerabilities in their Trusted OSes, receiving 10 new CVEs so far. Furthermore, compared to the baseline, SyzTrust has demonstrated significant improvements, including 66% higher code coverage, 651% higher state coverage, and 31% improved vulnerability-finding capability. We report all discovered new vulnerabilities to vendors and open source SyzTrust. Qinying Wang, Boyu Chang, Shouling Ji, Yuan Tian 0001, Xuhong Zhang 0002, Chenyang Lyu, Mathias Payer, Wenhai Wang, Raheem A. Beyah |
SP | 4 |
| 2024 | Remote Keylogging Attacks in Multi-user VR Applications
Zihao Su, Kunlin Cai, Reuben Beeler, Lukas Dresel, Allan Garcia, Ilya Grishchenko, Yuan Tian 0001, Christopher Krügel, Giovanni Vigna |
USENIX Security Symposium | 7 |
| 2024 | Privacy-Preserving Liveness Detection for Securing Smart Voice InterfacesabstractSmart speakers are widely used as the primary user interface in intelligent systems, including smart homes and industrial IoT. However, they are vulnerable to voice spoofing attacks which result in malicious command execution or privacy information leakage. Passive liveness detection, which thwarts voice spoofing via analyzing the collected audio rather than deploying sensors to distinguish between live-human and spoofing voices, has drawn increasing attention. But existing schemes either face performance degradation under environmental factor changes or require the user to keep fixed gestures, which limit their deployment in real-world scenarios. Besides, the space distributed property of smart speakers causes building a universal classifier for all involved users to be cumbersome and increases privacy leakage issues. To address the challenges mentioned above, we propose LIVEARRAY, an efficient, lightweight, and privacy-preserving passive liveness detection system. LIVEARRAY exploits a novel liveness feature, array fingerprint, which utilizes the microphone array inherently adopted by the smart speaker to improve the accuracy of liveness detection. LIVEARRAY's further employs the federated learning-based architecture to reduce the dataset collection overhead during classifier building and eliminate the potential privacy leakage during data transmission. Experimental results show that LIVEARRAY achieves an accuracy of 99.16%, which is superior to existing passive schemes Yan Meng 0001, Jiachun Li 0001, Haojin Zhu, Yuan Tian 0001, Jiming Chen 0001 |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2024 | One Bad Apple Spoils the Barrel: Understanding the Security Risks Introduced by Third-Party Components in IoT FirmwareabstractCurrently, the development of IoT firmware heavily depends on third-party components (TPCs) to improve development efficiency. Nevertheless, TPCs are not secure, and the vulnerabilities in TPCs will influence the security of IoT firmware. Existing works pay less attention to the vulnerabilities caused by TPCs, and we still lack a comprehensive understanding of the security impact of TPC vulnerability against firmware. To fill in the knowledge gap, we design and implementFirmSec, which leverages syntactical features and control-flow graph features to detect the TPCs in firmware, and then recognizes the corresponding vulnerabilities. Based onFirmSec, we present the first large-scale analysis of the security risks raised by TPCs on 34,136 firmware images. We successfully detect 584 TPCs and identify 128,757 vulnerabilities caused by 429 CVEs. Our in-depth analysis reveals the diversity of security risks in firmware and discovers some well-known vulnerabilities are still rooted in firmware. Besides, we explore the geographical distribution of vulnerable devices and confirm that the security situation of devices in different regions varies. Our analysis also indicates that vulnerabilities caused by TPCs in firmware keep growing with the boom of the IoT ecosystem. Further analysis shows 2,478 commercial firmware images have potentially violated GPL/AGPL licensing terms. Shouling Ji, Jiacheng Xu 0006, Yuan Tian 0001, Qiuyang Wei, Qinying Wang, Chenyang Lyu, Xuhong Zhang 0002, Changting Lin, JingZheng Wu, Raheem A. Beyah |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2023 | Retrieval Enhanced Data Augmentation for Question Answering on Privacy PoliciesabstractPrior studies in privacy policies frame the question answering (QA) task as identifying the most relevant text segment or a list of sentences from a policy document given a user query.Existing labeled datasets are heavily imbalanced (only a few relevant segments), limiting the QA performance in this domain.In this paper, we develop a data augmentation framework based on ensembling retriever models that captures the relevant text segments from unlabeled policy documents and expand the positive examples in the training set.In addition, to improve the diversity and quality of the augmented data, we leverage multiple pre-trained language models (LMs) and cascade them with noise reduction filter models.Using our augmented data on the PrivacyQA benchmark, we elevate the existing baseline by a large margin (10% F1) and achieve a new state-of-the-art F1 score of 50%.Our ablation studies provide further insights into the effectiveness of our approach. Md. Rizwan Parvez, Jianfeng Chi, Wasi Uddin Ahmad, Yuan Tian 0001, Kai-Wei Chang 0001 |
EACL | 4 |
| 2023 | Exploring Smart Commercial Building Occupants' Perceptions and Notification Preferences of Internet of Things Data Collection in the United StatesabstractData collection through the Internet of Things (IoT) devices, or smart devices, in commercial buildings enables possibilities for increased convenience and energy efficiency. However, such benefits face a large perceptual challenge when being implemented in practice, due to the different ways occupants working in the buildings understand and trust in the data collection. The semi-public, pervasive, and multi-modal nature of data collection in smart buildings points to the need to study occupants’ understanding of data collection and notification preferences. We conduct an online study with 492 participants in the US who report working in smart commercial buildings regarding: 1) awareness and perception of data collection in smart commercial buildings, 2) privacy notification preferences, and 3) potential factors for privacy notification preferences. We find that around half of the participants are not fully aware of the data collection and use practices of IoT even though they notice the presence of IoT devices and sensors. We also discover many misunderstandings around different data practices. The majority of participants want to be notified of data practices in smart buildings, and they prefer push notifications to passive ones such as websites or physical signs. Surprisingly, mobile app notification, despite being a popular channel for smart homes, is the least preferred method for smart commercial buildings. Tu Le, Alan Wang 0002, Yaxing Yao, Yuanyuan Feng, Arsalan Heydarian, Norman M. Sadeh, Yuan Tian 0001 |
EuroS&P | 7 |
| 2023 | MagFingerprint: A Magnetic Based Device Fingerprinting in Wireless Charging
Jiachun Li 0001, Yan Meng 0001, Guoxing Chen, Yuan Tian 0001, Haojin Zhu, Xuemin Shen |
INFOCOM | 5 |
| 2023 | CHKPLUG: Checking GDPR Compliance of WordPress Plugins via Cross-language Code Property Graph
Faysal Hossain Shezan, Zihao Su, Mingqing Kang, Nicholas Phair, Patrick William Thomas, Michelangelo van Dam, Yinzhi Cao, Yuan Tian 0001 |
NDSS | 8 |
| 2023 | What Distributions are Robust to Indiscriminate Poisoning Attacks for Linear Learners?abstractWe study indiscriminate poisoning for linear learners where an adversary injects a few crafted examples into the training data with the goal of forcing the induced model to incur higher test error. Inspired by the observation that linear learners on some datasets are able to resist the best known attacks even without any defenses, we further investigate whether datasets can be inherently robust to indiscriminate poisoning attacks for linear learners. For theoretical Gaussian distributions, we rigorously characterize the behavior of an optimal poisoning attack, defined as the poisoning strategy that attains the maximum risk of the induced model at a given poisoning budget. Our results prove that linear learners can indeed be robust to indiscriminate poisoning if the class-wise data distributions are well-separated with low variance and the size of the constraint set containing all permissible poisoning points is also small. These findings largely explain the drastic variation in empirical attack performance of the state-of-the-art poisoning attacks on linear learners across benchmark datasets, making an important initial step towards understanding the underlying reasons some learning tasks are vulnerable to data poisoning attacks. Fnu Suya, Xiao Zhang 0016, Yuan Tian 0001, David Evans 0001 |
NeurIPS | 3 |
| 2023 | Towards Usable Security Analysis Tools for Trigger-Action Programming
McKenna McCall, Eric Zeng 0001, Faysal Hossain Shezan, Mitchell Yang, Lujo Bauer, Abhishek Bichhawat, Camille Cobb, Limin Jia 0001, Yuan Tian 0001 |
SOUPS | 9 |
| 2023 | UVSCAN: Detecting Third-Party Component Usage Violations in IoT Firmware
Shouling Ji, Xuhong Zhang 0002, Yuan Tian 0001, Qinying Wang, Yuwen Pu, Chenyang Lyu, Raheem A. Beyah |
USENIX Security Symposium | 4 |
| 2023 | SenRev: Measurement of Personal Information Disclosure in Online Health CommunitiesabstractWith life style shifting during the pandemic, online health communities start to attract more users (including healthcare workers and patients) to discuss health-related questions. While such online platforms provide convenience to users, with health-related information shared broadly over text and images (e.g., X-Ray scans, photocopies of documents), they also raise questions regarding privacy. In this paper, we propose SenRev to systematically measure the leakages of sensitive information in those publicly available discussions. We use SenRev to analyze 1,894,900 multi-modal and multi-lingual data elements from four different online health communities. We find that sensitive data leakages are common; overall 1,324,064 (69.88%) pieces of evidence of data leakages are detected, with 23,587 (1.78%) of them involving identifiers and 1,300,477 (98.22%) involving quasi-identifiers. Surprisingly, leakages through medical images occur more frequently in the community of healthcare professionals compared with the other communities. Finally, based on our results, we discuss the potential directions for countermeasures. Faysal Hossain Shezan, Minjun Long, David Hasani, Gang Wang 0011, Yuan Tian 0001 |
Proc. Priv. Enhancing Technol. | 5 |
| 2022 | Towards Return Parity in Markov Decision ProcessesabstractAlgorithmic decisions made by machine learning models in high-stakes domains may have lasting impacts over time. However, naive applications of standard fairness criterion in static settings over temporal domains may lead to delayed and adverse effects. To understand the dynamics of performance disparity, we study a fairness problem in Markov decision processes (MDPs). Specifically, we propose return parity, a fairness notion that requires MDPs from different demographic groups that share the same state and action spaces to achieve approximately the same expected time-discounted rewards. We first provide a decomposition theorem for return disparity, which decomposes the return disparity of any two MDPs sharing the same state and action spaces into the distance between group-wise reward functions, the discrepancy of group policies, and the discrepancy between state visitation distributions induced by the group policies. Motivated by our decomposition theorem, we propose algorithms to mitigate return disparity via learning a shared group policy with state visitation distributional alignment using integral probability metrics. We conduct experiments to corroborate our results, showing that the proposed algorithm can successfully close the disparity gap while maintaining the performance of policies on two real-world recommender system benchmark datasets. Jianfeng Chi, Jian Shen 0003, Xinyi Dai, Weinan Zhang 0001, Yuan Tian 0001, Han Zhao 0002 |
AISTATS | 5 |
| 2022 | Cerberus: Query-driven Scalable Vulnerability Detection in OAuth Service Provider ImplementationsabstractOAuth protocols have been widely adopted to simplify user authentication and service authorization for third-party applications. However, little effort has been devoted to automatically checking the security of the libraries that service providers widely use. In this paper, we formalize the OAuth specifications and security best practices, and design Cerberus, an automated static analyzer, to find logical flaws and identify vulnerabilities in the implementation of OAuth service provider libraries. To efficiently detect security violations in a large codebase of service provider implementation, Cerberus employs a query-driven algorithm for answering queries about OAuth specifications. We demonstrate the effectiveness of Cerberus by evaluating it on datasets of popular OAuth libraries with millions of downloads. Among these high-profile libraries, Cerberus has identified 47 vulnerabilities from ten classes of logical flaws, 24 of which were previously unknown. We got acknowledged by the developers of eight libraries and had three accepted CVEs. Tamjid Al Rahat, Yu Feng 0001, Yuan Tian 0001 |
CCS | 3 |
| 2022 | A large-scale empirical analysis of the vulnerabilities introduced by third-party components in IoT firmwareabstractAs the core of IoT devices, firmware is undoubtedly vital. Currently, the development of IoT firmware heavily depends on third-party components (TPCs), which significantly improves the development efficiency and reduces the cost. Nevertheless, TPCs are not secure, and the vulnerabilities in TPCs will turn back influence the security of IoT firmware. Currently, existing works pay less attention to the vulnerabilities caused by TPCs, and we still lack a comprehensive understanding of the security impact of TPC vulnerability against firmware. To fill in the knowledge gap, we design and implement FirmSec, which leverages syntactical features and control-flow graph features to detect the TPCs at version-level in firmware, and then recognizes the corresponding vulnerabilities. Based on FirmSec, we present the first large-scale analysis of the usage of TPCs and the corresponding vulnerabilities in firmware. More specifically, we perform an analysis on 34,136 firmware images, including 11,086 publicly accessible firmware images, and 23,050 private firmware images from TSmart. We successfully detect 584 TPCs and identify 128,757 vulnerabilities caused by 429 CVEs. Our in-depth analysis reveals the diversity of security issues for different kinds of firmware from various vendors, and discovers some well-known vulnerabilities are still deeply rooted in many firmware images. We also find that the TPCs used in firmware have fallen behind by five years on average. Besides, we explore the geographical distribution of vulnerable devices, and confirm the security situation of devices in several regions, e.g., South Korea and China, is more severe than in other regions. Further analysis shows 2,478 commercial firmware images have potentially violated GPL/AGPL licensing terms. Shouling Ji, Jiacheng Xu 0006, Yuan Tian 0001, Qiuyang Wei, Qinying Wang, Chenyang Lyu, Xuhong Zhang 0002, Changting Lin, JingZheng Wu, Raheem A. Beyah |
ISSTA | 4 |
| 2022 | Your Microphone Array Retains Your Identity: A Robust Voice Liveness Detection System for Smart Speakers
Yan Meng 0001, Jiachun Li 0001, Matthew Pillari, Arjun Deopujari, Liam Brennan, Hafsah Shamsie, Haojin Zhu, Yuan Tian 0001 |
USENIX Security Symposium | 8 |
| 2022 | SkillBot: Identifying Risky Content for Children in Alexa SkillsabstractMany households include children who use voice personal assistants (VPA) such as Amazon Alexa. Children benefit from the rich functionalities of VPAs and third-party apps but are also exposed to new risks in the VPA ecosystem. In this article, we first investigate “risky” child-directed voice apps that contain inappropriate content or ask for personal information through voice interactions. We build SkillBot—a natural language processing-based system to automatically interact with VPA apps and analyze the resulting conversations. We find 28 risky child-directed apps and maintain a growing dataset of 31,966 non-overlapping app behaviors collected from 3,434 Alexa apps. Our findings suggest that although child-directed VPA apps are subject to stricter policy requirements and more intensive vetting, children remain vulnerable to inappropriate content and privacy violations. We then conduct a user study showing that parents are concerned about the identified risky apps. Many parents do not believe that these apps are available and designed for families/kids, although these apps are actually published in Amazon’s “Kids” product category. We also find that parents often neglect basic precautions, such as enabling parental controls on Alexa devices. Finally, we identify a novel risk in the VPA ecosystem: confounding utterances or voice commands shared by multiple apps that may cause a user to interact with a different app than intended. We identify 4,487 confounding utterances, including 581 shared by child-directed and non-child-directed apps. We find that 27% of these confounding utterances prioritize invoking a non-child-directed app over a child-directed app. This indicates that children are at real risk of accidentally invoking non-child-directed apps due to confounding utterances. Tu Le, Danny Yuxing Huang, Noah J. Apthorpe, Yuan Tian 0001 |
ACM Trans. Internet Techn. | 4 |
| 2021 | Curse or Redemption? How Data Heterogeneity Affects the Robustness of Federated LearningabstractData heterogeneity has been identified as one of the key features in federated learning but often overlooked in the lens of robustness to adversarial attacks. This paper focuses on characterizing and understanding its impact on backdooring attacks in federated learning through comprehensive experiments using synthetic and the LEAF benchmarks. The initial impression driven by our experimental results suggests that data heterogeneity is the dominant factor in the effectiveness of attacks and it may be a redemption for defending against backdooring as it makes the attack less efficient, more challenging to design effective attack strategies, and the attack result also becomes less predictable. However, with further investigations, we found data heterogeneity is more of a curse than a redemption as the attack effectiveness can be significantly boosted by simply adjusting the client-side backdooring timing. More importantly, data heterogeneity may result in overfitting at the local training of benign clients, which can be utilized by attackers to disguise themselves and fool skewed-feature based defenses. In addition, effective attack strategies can be made by adjusting attack data distribution. Finally, we discuss the potential directions of defending the curses brought by data heterogeneity. The results and lessons learned from our extensive experiments and analysis offer new insights for designing robust federated learning methods and systems. Syed Zawad, Ali Anwar 0001, Yi Zhou 0015, Nathalie Baracaldo, Yuan Tian 0001, Feng Yan 0001 |
AAAI | 7 |
| 2021 | Intent Classification and Slot Filling for Privacy PoliciesabstractWasi Ahmad, Jianfeng Chi, Tu Le, Thomas Norton, Yuan Tian, Kai-Wei Chang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Wasi Uddin Ahmad, Jianfeng Chi, Tu Le, Thomas Norton, Yuan Tian 0001, Kai-Wei Chang 0001 |
ACL/IJCNLP (1) | 5 |
| 2021 | Understanding and Mitigating Accuracy Disparity in RegressionabstractWith the widespread deployment of large-scale prediction systems in high-stakes domains, e.g., face recognition, criminal justice, etc., disparity on prediction accuracy between different demographic subgroups has called for fundamental understanding on the source of such disparity and algorithmic intervention to mitigate it. In this paper, we study the accuracy disparity problem in regression. To begin with, we first propose an error decomposition theorem, which decomposes the accuracy disparity into the distance between marginal label distributions and the distance between conditional representations, to help explain why such accuracy disparity appears in practice. Motivated by this error decomposition and the general idea of distribution alignment with statistical distances, we then propose an algorithm to reduce this disparity, and analyze its game-theoretic optima of the proposed objective functions. To corroborate our theoretical findings, we also conduct experiments on five benchmark datasets. The experimental results suggest that our proposed algorithms can effectively mitigate accuracy disparity while maintaining the predictive power of the regression models. Jianfeng Chi, Yuan Tian 0001, Geoffrey J. Gordon, Han Zhao 0002 |
ICML | 2 |
| 2021 | Model-Targeted Poisoning Attacks with Provable ConvergenceabstractIn a poisoning attack, an adversary who controls a small fraction of the training data attempts to select that data, so a model is induced that misbehaves in a particular way. We consider poisoning attacks against convex machine learning models and propose an efficient poisoning attack designed to induce a model specified by the adversary. Unlike previous model-targeted poisoning attacks, our attack comes with provable convergence to any attainable target model. We also provide a lower bound on the minimum number of poisoning points needed to achieve a given target model. Our method uses online convex optimization and finds poisoning points incrementally. This provides more flexibility than previous attacks which require an a priori assumption about the number of poisoning points. Our attack is the first model-targeted poisoning attack that provides provable convergence for convex models. In our experiments, it either exceeds or matches state-of-the-art attacks in terms of attack success rate and distance to the target model. Fnu Suya, Saeed Mahloujifar, Anshuman Suri, David Evans 0001, Yuan Tian 0001 |
ICML | 5 |
| 2021 | Hardware/Software Security Patches for the Internet of ThingsabstractWith the rapid development of the Internet of Things (IoT), there are billions of interacting devices and applications. With so many devices and applications, one of the most critical challenges is how to provide security. Traditional software-based defenses will not be enough to protect the security of IoT because of the attack surfaces derived from the physical environment. For example, an attacker can physically re-point a surveillance camera, can move a smart device to another location, can send a sound signal to influence an accelerometer, can cause wireless jamming, etc. We propose to create "smart buttons," and collections of them called "smart blankets" as hardware/software (HW/SW) security patches rather than software-only patches. These fixes operate similarly to software patches, but because of the hardware added, these new patches can better support against physical world attacks. While this paper primarily presents a vision for HW/SW patches, solutions are implemented and shown for two classes of attacks involving cameras and robots. Open questions are also discussed. John A. Stankovic, Tu Le, Abdeltawab M. Hendawi, Yuan Tian 0001 |
SMARTCOMP | 4 |
| 2021 | CryptGPU: Fast Privacy-Preserving Machine Learning on the GPUabstractWe introduce CryptGPU, a system for privacy-preserving machine learning that implements all operations on the GPU (graphics processing unit). Just as GPUs played a pivotal role in the success of modern deep learning, they are also essential for realizing scalable privacy-preserving deep learning. In this work, we start by introducing a new interface to losslessly embed cryptographic operations over secret-shared values (in a discrete domain) into floating-point operations that can be processed by highly-optimized CUDA kernels for linear algebra. We then identify a sequence of "GPU-friendly" cryptographic protocols to enable privacy-preserving evaluation of both linear and non-linear operations on the GPU. Our microbenchmarks indicate that our private GPU-based convolution protocol is over 150× faster than the analogous CPU-based protocol; for non-linear operations like the ReLU activation function, our GPU-based protocol is around 10× faster than its CPU analog. With CryptGPU, we support private inference and training on convolutional neural networks with over 60 million parameters as well as handle large datasets like ImageNet. Compared to the previous state-of-the-art, our protocols achieve a 2× to 8× improvement in private inference for large networks and datasets. For private training, we achieve a 6× to 36× improvement over prior state-of-the-art. Our work not only showcases the viability of performing secure multiparty computation (MPC) entirely on the GPU to newly enable fast privacy-preserving machine learning, but also highlights the importance of designing new MPC primitives that can take full advantage of the GPU’s computing capabilities. Sijun Tan, Brian Knott, Yuan Tian 0001, David J. Wu 0001 |
SP | 3 |
| 2021 | MPInspector: A Systematic and Automatic Approach for Evaluating the Security of IoT Messaging Protocols
Qinying Wang, Shouling Ji, Yuan Tian 0001, Xuhong Zhang 0002, Yuhong Kan, Zhaowei Lin, Changting Lin, Shuiguang Deng, Alex X. Liu, Raheem A. Beyah |
USENIX Security Symposium | 3 |
| 2020 | TKPERM: Cross-platform Permission Knowledge Transfer to Detect Overprivileged Third-party Applications
Faysal Hossain Shezan, Kaiming Cheng, Yinzhi Cao, Yuan Tian 0001 |
NDSS | 5 |
| 2020 | Trade-offs and Guarantees of Adversarial Representation Learning for Information ObfuscationabstractCrowdsourced data used in machine learning services might carry sensitive information about attributes that users do not want to share. Various methods have been proposed to minimize the potential information leakage of sensitive attributes while maximizing the task accuracy. However, little is known about the theory behind these methods. In light of this gap, we develop a novel theoretical framework for attribute obfuscation. Under our framework, we propose a minimax optimization formulation to protect the given attribute and analyze its inference guarantees against worst-case adversaries. Meanwhile, there is a tension between minimizing information leakage and maximizing task accuracy. To understand this, we prove an information-theoretic lower bound to precisely characterize the fundamental trade-off between accuracy and information leakage. We conduct experiments on two real-world datasets to corroborate the inference guarantees and validate the inherent trade-offs therein. Our results indicate that, among several alternatives, the adversarial learning approach achieves the best trade-off in terms of attribute obfuscation and accuracy maximization. Han Zhao 0002, Jianfeng Chi, Yuan Tian 0001, Geoffrey J. Gordon |
NeurIPS | 3 |
| 2020 | Hybrid Batch Attacks: Finding Black-box Adversarial Examples with Limited Queries
Fnu Suya, Jianfeng Chi, David Evans 0001, Yuan Tian 0001 |
USENIX Security Symposium | 4 |
| 2020 | iOS, Your OS, Everybody's OS: Vetting and Analyzing Network Services of iOS Applications
Zhushou Tang, Minhui Xue 0001, Yuan Tian 0001, Sen Chen 0001, Muhammad Ikram 0001, Tielei Wang, Haojin Zhu |
USENIX Security Symposium | 4 |
| 2020 | Evaluating the Dedicated Short-range Communication for Connected Vehicles against Network Security Attacks
Tu Le, Ingy Elsayed-Aly, Weizhao Jin, Seunghan Ryu, Guy Verrier, Tamjid Al Rahat, Yuan Tian 0001 |
VEHITS | 8 |
| 2020 | Read Between the Lines: An Empirical Measurement of Sensitive Applications of Voice Personal Assistant SystemsabstractVoice Personal Assistant (VPA) systems such as Amazon Alexa and Google Home have been used by tens of millions of households. Recent work demonstrated proof-of-concept attacks against their voice interface to invoke unintended applications or operations. However, there is still a lack of empirical understanding of what type of third-party applications that VPA systems support, and what consequences these attacks may cause. In this paper, we perform an empirical analysis of the third-party applications of Amazon Alexa and Google Home to systematically assess the attack surfaces. A key methodology is to characterize a given application by classifying the sensitive voice commands it accepts. We develop a natural language processing tool that classifies a given voice command from two dimensions: (1) whether the voice command is designed to insert action or retrieve information; (2) whether the command is sensitive or nonsensitive. The tool combines a deep neural network and a keyword-based model, and uses Active Learning to reduce the manual labeling effort. The sensitivity classification is based on a user study (N=404) where we measure the perceived sensitivity of voice commands. A ground-truth evaluation shows that our tool achieves over 95% of accuracy for both types of classifications. We apply this tool to analyze 77,957 Amazon Alexa applications and 4,813 Google Home applications (198,199 voice commands from Amazon Alexa, 13,644 voice commands from Google Home) over two years (2018-2019). In total, we identify 19,263 sensitive “action injection” commands and 5,352 sensitive “information retrieval” commands. These commands are from 4,596 applications (5.55% out of all applications), most of which belong to the “smart home” category. While the percentage of sensitive applications is small, we show the percentage is increasing over time from 2018 to 2019. Faysal Hossain Shezan, Hang Hu 0002, Gang Wang 0011, Yuan Tian 0001 |
WWW | 5 |
| 2019 | StopGuessing: Using Guessed Passwords to Thwart Online GuessingabstractPractitioners who seek to defend password-protected resources from online guessing attacks will find a shortage of tooling and techniques to help them. Little research suggests anything beyond blocking or throttling traffic from IP addresses sending suspicious traffic; counting failed authentication requests, or some variant, is often the sole feature used to determine suspicion. In this paper we show that several other features can greatly help distinguishing benign and attack traffic. First, we increase the penalties for clients responsible for fail events involving passwords frequently-guessed by attackers. Second, we reduce the threshold (and thus protect better) for accounts with weak passwords. Third, we detect, and are more forgiving of, login failures caused by users mistyping their passwords. Most importantly, we achieve all of these goals without needing any marker that indicates weak accounts, changing the format in which passwords are stored (i.e. we do not store passwords plaintext or in any recoverable form), or storing any information that might be harmful if leaked. We present an open-source implementation of this system and demonstrate its improvement over simpler blocking strategies in various simulated scenarios. Stuart E. Schechter, Yuan Tian 0001, Cormac Herley |
EuroS&P | 2 |
| 2019 | OAUTHLINT: An Empirical Study on OAuth Bugs in Android ApplicationsabstractMobile developers use OAuth APIs to implement Single-Sign-On services. However, the OAuth protocol was originally designed for the authorization for third-party websites not to authenticate users in third-party mobile apps. As a result, it is challenging for developers to correctly implement mobile OAuth securely. These vulnerabilities due to the misunderstanding of OAuth and inexperience of developers could lead to data leakage and account breach. In this paper, we perform an empirical study on the usage of OAuth APIs in Android applications and their security implications. In particular, we develop OAUTHLINT, that incorporates a query-driven static analysis to automatically check programs on the Google Play marketplace. OAUTHLINT takes as input an anti-protocol that encodes a vulnerable pattern extracted from the OAuth specifications and a program P. Our tool then generates a counter-example if the anti-protocol can match a trace of Ps possible executions. To evaluate the effectiveness of our approach, we perform a systematic study on 600+ popular apps which have more than 10 millions of downloads. The evaluation shows that 101 (32%) out of 316 applications that use OAuth APIs make at least one security mistake. Tamjid Al Rahat, Yu Feng 0001, Yuan Tian 0001 |
ASE | 3 |
| 2019 | Demystifying Hidden Privacy Settings in Mobile AppsabstractMobile apps include privacy settings that allow their users to configure how their data should be shared. These settings, however, are often hard to locate and hard to understand by the users, even in popular apps, such as Facebook. More seriously, they are often set to share user data by default, exposing her privacy without proper consent. In this paper, we report the first systematic study on the problem, which is made possible through an in-depth analysis of user perception of the privacy settings. More specifically, we first conduct two user studies (involving nearly one thousand users) to understand privacy settings from the user's perspective, and identify these hard-to-find settings. Then we select 14 features that uniquely characterize such hidden privacy settings and utilize a novel technique called semantics- based UI tracing to extract them from a given app. On top of these features, a classifier is trained to automatically discover the hidden privacy settings, which together with other innovations, has been implemented into a tool called Hound. Over our labeled data set, the tool achieves an accuracy of 93.54%. Further running it on 100,000 latest apps from both Google Play and third-party markets, we find that over a third (36.29%) of the privacy settings identified from these apps are “hidden”. Looking into these settings, we observe that they become hard to discover and hard to understand primarily due to the problematic categorization on the apps' user interfaces and/or confusing descriptions. Further importantly, though more privacy options have been offered to the user over time, also discovered is the persistence of their usability issue, which becomes even more serious, e.g., originally easy-to-find settings now harder to locate. And among all such hidden privacy settings, 82.16% are set to leak user privacy by default. We provide suggestions for improving the usability of these privacy settings at the end of our study. Yi Chen 0024, Mingming Zha 0001, Nan Zhang 0018, Dandan Xu, Xuan Feng 0005, Kan Yuan, Fnu Suya, Yuan Tian 0001, Kai Chen 0012, XiaoFeng Wang 0001 |
IEEE Symposium on Security and Privacy | 9 |
| 2019 | Dangerous Skills: Understanding and Mitigating Security Risks of Voice-Controlled Third-Party Functions on Virtual Personal Assistant SystemsabstractVirtual personal assistants (VPA) (e.g., Amazon Alexa and Google Assistant) today mostly rely on the voice channel to communicate with their users, which however is known to be vulnerable, lacking proper authentication (from the user to the VPA). A new authentication challenge, from the VPA service to the user, has emerged with the rapid growth of the VPA ecosystem, which allows a third party to publish a function (called skill) for the service and therefore can be exploited to spread malicious skills to a large audience during their interactions with smart speakers like Amazon Echo and Google Home. In this paper, we report a study that concludes such remote, large-scale attacks are indeed realistic. We discovered two new attacks: voice squatting in which the adversary exploits the way a skill is invoked (e.g., ``open capital one''), using a malicious skill with a similarly pronounced name (e.g., ``capital won'') or a paraphrased name (e.g., ``capital one please'') to hijack the voice command meant for a legitimate skill (e.g., ``capital one''), and voice masquerading in which a malicious skill impersonates the VPA service or a legitimate skill during the user's conversation with the service to steal her personal information. These attacks aim at the way VPAs work or the user's misconceptions about their functionalities, and are found to pose a realistic threat by our experiments (including user studies and real-world deployments) on Amazon Echo and Google Home. The significance of our findings has already been acknowledged by Amazon and Google, and further evidenced by the risky skills found on Alexa and Google markets by the new squatting detector we built. We further developed a technique that automatically captures an ongoing masquerading attack and demonstrated its efficacy. Nan Zhang 0018, Xianghang Mi, Xuan Feng 0005, XiaoFeng Wang 0001, Yuan Tian 0001, Feng Qian 0001 |
IEEE Symposium on Security and Privacy | 5 |
| 2019 | Birthday, Name and Bifacial-security: Understanding Passwords of Chinese Web Users
Ding Wang 0002, Ping Wang 0003, Debiao He, Yuan Tian 0001 |
USENIX Security Symposium | 4 |
| 2017 | IVD: Automatic Learning and Enforcement of Authorization Rules in Online Social NetworksabstractAuthorization bugs, when present in online social networks, are usually caused by missing or incorrect authorization checks and can allow attackers to bypass the online social network's protections. Unfortunately, there is no practical way to fully guarantee that an authorization bug will never be introduced-even with good engineering practices-as a web application and its data model become more complex. Unlike other web application vulnerabilities such as XSS and CSRF, there is no practical general solution to prevent missing or incorrect authorization checks. In this paper we propose Invariant Detector (IVD), a defense-in-depth system that automatically learns authorization rules from normal data manipulation patterns and distills them into likely invariants. These invariants, usually learned during the testing or pre-release stages of new features, are then used to block any requests that may attempt to exploit bugs in the social network's authorization logic. IVD acts as an additional layer of defense, working behind the scenes, complementary to privacy frameworks and testing. We have designed and implemented IVD to handle the unique challenges posed by modern online social networks. IVD is currently running at Facebook, where it infers and evaluates daily more than 200,000 invariants from a sample of roughly 500 million client requests, and checks the resulting invariants every second against millions of writes made to a graph database containing trillions of entities. Thus far IVD has detected several high impact authorization bugs and has successfully blocked attempts to exploit them before code fixes were deployed. Paul Marinescu, Chad Parry, Marjori Pomarole, Yuan Tian 0001, Patrick Tague, Ioannis Papagiannis |
IEEE Symposium on Security and Privacy | 4 |
| 2017 | SmartAuth: User-Centered Authorization for the Internet of Things
Yuan Tian 0001, Nan Zhang 0018, Yue-Hsun Lin, XiaoFeng Wang 0001, Blase Ur, Xianzheng Guo, Patrick Tague |
USENIX Security Symposium | 1 |
| 2016 | Swords and shields: a study of mobile game hacks and existing defenses
Yuan Tian 0001, Eric Yawei Chen, Shuo Chen 0001, Xiao Wang 0040, Patrick Tague |
ACSAC | 1 |
| 2015 | The Activity Platform
Helen J. Wang, Alexander Moshchuk, Michael Gamon, Shamsi T. Iqbal, Eli T. Brown, Ashish Kapoor, Christopher Meek, Eric Yawei Chen, Yuan Tian 0001, Jaime Teevan, Mary Czerwinski, Susan T. Dumais |
HotOS | 9 |
| 2015 | Run-time Monitoring and Formal Analysis of Information Flows in Chromium
Lujo Bauer, Shaoying Cai, Limin Jia 0001, Timothy Passaro, Michael Stroucken, Yuan Tian 0001 |
NDSS | 6 |
| 2014 | OAuth Demystified for Mobile Application DevelopersabstractOAuth has become a highly influential protocol due to its swift and wide adoption in the industry. The initial objective of the protocol was specific: it serves the authorization needs for websites. What motivates our work is the realization that the protocol has been significantly re-purposed and re-targeted over the years: (1) all major identity providers, e.g., Facebook, Google and Microsoft, have re-purposed OAuth for user authentication; (2) developers have re-targeted OAuth to the mobile platforms, in addition to the traditional web platform. Therefore, we believe that it is necessary and timely to conduct an in-depth study to demystify OAuth for mobile application developers. Our work consists of two pillars: (1) an in-house study of the OAuth protocol documentation that aims to identify what might be ambiguous or unspecified for mobile developers; (2) a field-study of over 600 popular mobile applications that highlights how well developers fulfill the authentication and authorization goals in practice. The result is really worrisome: among the 149 applications that use OAuth, 89 of them (59.7%) were incorrectly implemented and thus vulnerable. In the paper, we pinpoint the key portions in each OAuth protocol flow that are security critical, but are confusing or unspecified for mobile application developers. We then show several representative cases to concretely explain how real implementations fell into these pitfalls. Our findings have been communicated to vendors of the vulnerable applications. Most vendors positively confirmed the issues, and some have applied fixes. We summarize lessons learned from the study, hoping to provoke further thoughts about clear guidelines for OAuth usage in mobile applications. Eric Yawei Chen, Yutong Pei, Shuo Chen 0001, Yuan Tian 0001, Robert Kotcher, Patrick Tague |
CCS | 4 |
| 2014 | All Your Screens Are Belong to Us: Attacks Exploiting the HTML5 Screen Sharing APIabstractHTML5 changes many aspects in the browser world by introducing numerous new concepts, in particular, the new HTML5 screen sharing API impacts the security implications of browsers tremendously. One of the core assumptions on which browser security is built is that there is no cross-origin feedback loop from the client to the server. However, the screen sharing API allows creating a cross-origin feedback loop. Consequently, websites will potentially be able to see all visible content from the user's screen, irrespective of its origin. This cross-origin feedback loop, when combined with human vision limitations, can introduce new vulnerabilities. An attacker can capture sensitive information from victim's screen using the new API without the consensus of the victim. We investigate the security implications of the screen sharing API and discuss how existing defenses against traditional web attacks fail during screen sharing. We show that several attacks are possible with the help of the screen sharing API: cross-site request forgery, history sniffing, and information stealing. We discuss how popular websites such as Amazon and Wells Fargo can be attacked using this API and demonstrate the consequences of the attacks such as economic losses, compromised account and information disclosure. The objective of this paper is to present the attacks using the screen sharing API, analyze the fundamental cause and motivate potential defenses to design a more secure screen sharing API. Yuan Tian 0001, Ying Chuan Liu, Amar Bhosale, Lin-Shung Huang, Patrick Tague, Collin Jackson |
IEEE Symposium on Security and Privacy | 1 |
| 2014 | PrivateDroid: Private Browsing Mode for AndroidabstractPrivate browsing mode is a privacy feature adopted by many modern computer browsers. With the increased use of mobile devices and escalating privacy concerns for mobile users, browser applications on mobile devices have also started incorporating private browsing mode. Even so, the use of private browsing mode is limited to the browser applications and cannot be applied directly on other third-party mobile applications. In this paper, we propose Private Droid, which provides a private browsing mode for third-party applications on the Android platform. First, we discuss three possible approaches of implementing mobile private browsing mode: code instrumentation, an extra sandbox, and a Linux container approach. Then, we implement Private Droid, which creates a new sandbox for every application in private mode and destroys the sandbox once the application is closed. After that, we evaluate usability, efficiency and security of the system with 25 popular Android applications. Our design considerations, implementation details, evaluation results, and challenges lay a foundation of private browsing mode on mobile platforms. Su Mon Kywe, Christopher B. Landis, Yutong Pei, Justin Satterfield, Yuan Tian 0001, Patrick Tague |
TrustCom | 5 |