VLDB 2026 Research / reviewers in the wild / expert
Yuanshun Yao
dblp:186/1486
· DBLP profile ↗
18ranked-venue papers
4as first author
10since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 1 first-author · 9 since 2021Security and privacy · 5 · 2 first-author · 1 since 2021Computer networks · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 2Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
11 papers |
Trustworthy machine learning · 46% Language models and text generation · 26% Efficient and distributed learning · 10% | |
| Network and information security
7 papers |
Security and privacy of machine learning · 49% Digital forensics and information hiding · 26% Privacy and data protection · 20% | |
| Databases, data mining, and information retrieval
2 papers |
Data mining · 72% Machine learning and data management · 28% | |
| Computer networks
2 papers |
Wireless sensing and localization · 100% |
Topics — the 30 heaviest of 39, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
fairness |
2.2 | 3 | 2024 | Fairness without Harm: An Influence-Guided Active Sampling Approach · NeurIPS 2024 Fair Classifiers that Abstain without Harm · ICLR 2024 Weak Proxies are Sufficient and Preferable for Fairness with Missing Sensitive Attributes · ICML 2023 |
Machine learning › Trustworthy machine learning › fairness
group fairness |
1.5 | 2 | 2024 | Fairness without Harm: An Influence-Guided Active Sampling Approach · NeurIPS 2024 Fair Classifiers that Abstain without Harm · ICLR 2024 |
Security and privacy of machine learning › adversarial attack
backdoor attack |
1.3 | 3 | 2021 | Backdoor Attacks Against Deep Learning Systems in the Physical World · CVPR 2021 Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural Networks · IEEE Symposium on Security and Privacy 2019 Latent Backdoor Attacks on Deep Neural Networks · CCS 2019 |
Security and privacy of machine learning › adversarial attack › backdoor attack
backdoor defense |
0.9 | 2 | 2021 | Backdoor Attacks Against Deep Learning Systems in the Physical World · CVPR 2021 Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural Networks · IEEE Symposium on Security and Privacy 2019 |
Machine learning › Reinforcement learning
actor-critic methods |
0.9 | 1 | 2025 | ACC-Collab: An Actor-Critic Approach to Multi-Agent LLM Collaboration · ICLR 2025 |
Knowledge, reasoning and agents › Multi-agent systems
multi-agent collaboration |
0.9 | 1 | 2025 | ACC-Collab: An Actor-Critic Approach to Multi-Agent LLM Collaboration · ICLR 2025 |
Digital forensics and information hiding › watermarking
text watermarking |
0.9 | 1 | 2025 | Robust Multi-bit Text Watermark with LLM-based Paraphrasers · ICML 2025 |
Digital forensics and information hiding
watermarking |
0.9 | 1 | 2025 | Robust Multi-bit Text Watermark with LLM-based Paraphrasers · ICML 2025 |
Natural language and speech › Language models and text generation
alignment |
0.8 | 1 | 2024 | Large Language Model Unlearning · NeurIPS 2024 |
Machine learning › Trustworthy machine learning › fairness › fairness trade-off
fairness-accuracy trade-off |
0.8 | 1 | 2024 | Fairness without Harm: An Influence-Guided Active Sampling Approach · NeurIPS 2024 |
Natural language and speech › Language models and text generation
hallucination mitigation |
0.8 | 1 | 2024 | Large Language Model Unlearning · NeurIPS 2024 |
Natural language and speech › Language models and text generation
large language model |
0.8 | 1 | 2024 | Large Language Model Unlearning · NeurIPS 2024 |
Machine learning › Trustworthy machine learning
machine unlearning |
0.8 | 1 | 2024 | Large Language Model Unlearning · NeurIPS 2024 |
Machine learning › Trustworthy machine learning
robustness |
0.8 | 1 | 2024 | Large Language Model Unlearning · NeurIPS 2024 |
Machine learning › Efficient and distributed learning › federated learning
federated evaluation |
0.7 | 1 | 2023 | DPAUC: Differentially Private AUC Computation in Federated Learning · AAAI 2023 |
Machine learning › Efficient and distributed learning
federated learning |
0.7 | 1 | 2023 | DPAUC: Differentially Private AUC Computation in Federated Learning · AAAI 2023 |
Privacy and data protection
differential privacy |
0.7 | 1 | 2023 | DPAUC: Differentially Private AUC Computation in Federated Learning · AAAI 2023 |
Privacy and data protection › differential privacy › relaxed differential privacy
label differential privacy |
0.7 | 1 | 2023 | DPAUC: Differentially Private AUC Computation in Federated Learning · AAAI 2023 |
Security and privacy of machine learning › adversarial attack › backdoor attack
physical backdoor attack |
0.5 | 1 | 2021 | Backdoor Attacks Against Deep Learning Systems in the Physical World · CVPR 2021 |
Security and privacy of machine learning › adversarial attack › backdoor attack › backdoor defense
backdoor detection |
0.4 | 1 | 2019 | Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural Networks · IEEE Symposium on Security and Privacy 2019 |
Security and privacy of machine learning
adversarial attack |
0.3 | 1 | 2018 | With Great Training Comes Great Vulnerability: Practical Attacks against Transfer Learning · USENIX Security Symposium 2018 |
Robotics › Robot navigation and mapping › obstacle avoidance
collision-free navigation |
0.3 | 1 | 2017 | Object Recognition and Navigation using a Single Networking Device · MobiSys 2017 |
Data mining
clustering |
0.3 | 1 | 2017 | Identifying Value in Crowdsourced Wireless Signal Measurements · WWW 2017 |
Data mining › clustering
feature clustering |
0.3 | 1 | 2017 | Identifying Value in Crowdsourced Wireless Signal Measurements · WWW 2017 |
Wireless sensing and localization › network localization
basestation localization |
0.3 | 1 | 2017 | Identifying Value in Crowdsourced Wireless Signal Measurements · WWW 2017 |
Wireless sensing and localization
RF imaging |
0.3 | 1 | 2017 | Object Recognition and Navigation using a Single Networking Device · MobiSys 2017 |
Performance modeling and evaluation
benchmarking |
0.3 | 1 | 2017 | Complexity vs. performance: empirical analysis of machine learning as a service · Internet Measurement Conference 2017 |
Natural language and speech › Language models and text generation
LLM agents |
0.3 | 1 | 2025 | ACC-Collab: An Actor-Critic Approach to Multi-Agent LLM Collaboration · ICLR 2025 |
Machine learning and data management
active learning |
0.2 | 1 | 2024 | Fairness without Harm: An Influence-Guided Active Sampling Approach · NeurIPS 2024 |
Computer vision › Face, body and person analysis
face recognition |
0.1 | 1 | 2021 | Backdoor Attacks Against Deep Learning Systems in the Physical World · CVPR 2021 |
Methods — techniques the papers use, named apart from their topics
text classification · 1.7fine-tuning · 1.7unlearning · 1.5influence-guided active sampling · 1.5generalization bound analysis · 1.5large language model · 0.9actor-critic · 0.9surrogate model · 0.8reinforcement learning from human feedback · 0.8integer programming · 0.8differential privacy · 0.7AUC computation · 0.7unsupervised learning · 0.6supervised learning · 0.6physical trigger · 0.5empirical study · 0.5transfer learning · 0.4neuron pruning · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ACC-Collab: An Actor-Critic Approach to Multi-Agent LLM CollaborationabstractLarge language models (LLMs) have demonstrated a remarkable ability to serve as general-purpose tools for various language-based tasks.
Recent works have demonstrated that the efficacy of such models can be improved through iterative dialog between multiple models.
While these
paradigms show promise
in
improving model efficacy, most works in this area treat collaboration as an emergent behavior, rather than a learned behavior.
In doing so, current multi-agent frameworks rely on collaborative behaviors to have been sufficiently trained into off-the-shelf models.
To address this limitation, we propose ACC-Collab, an **A**ctor-**C**riti**c** based learning framework to produce a two-agent team (an actor-agent and a critic-agent) specialized in collaboration.
We demonstrate that ACC-Collab outperforms SotA multi-agent techniques on a wide array of benchmarks. Andrew Estornell, Jean-Francois Ton, Yuanshun Yao, Yang Liu 0018 |
ICLR | 3 |
| 2025 | Robust Multi-bit Text Watermark with LLM-based ParaphrasersabstractWe propose an imperceptible multi-bit text watermark embedded by paraphrasing with LLMs. We fine-tune a pair of LLM paraphrasers that are designed to behave differently so that their paraphrasing difference reflected in the text semantics can be identified by a trained decoder. To embed our multi-bit watermark, we use two paraphrasers alternatively to encode the pre-defined binary code at the sentence level. Then we use a text classifier as the decoder to decode each bit of the watermark. Through extensive experiments, we show that our watermarks can achieve over 99.99% detection AUC with small (1.1B) text paraphrasers while keeping the semantic information of the original sentence. More importantly, our pipeline is robust under word substitution and sentence paraphrasing perturbations and generalizes well to out-of-distributional data. We also show the stealthiness of our watermark with LLM-based evaluation. Jinghan Jia, Yuanshun Yao, Hang Li 0001 |
ICML | 3 |
| 2024 | Fair Classifiers that Abstain without HarmabstractIn critical applications, it is vital for classifiers to defer decision-making to humans. We propose a post-hoc method that makes existing classifiers selectively abstain from predicting certain samples. Our abstaining classifier is incentivized to maintain the original accuracy for each sub-population (i.e. no harm) while achieving a set of group fairness definitions to a user specified degree. To this end, we design an Integer Programming (IP) procedure that assigns abstention decisions for each training sample to satisfy a set of constraints. To generalize the abstaining decisions to test samples, we then train a surrogate model to learn the abstaining decisions based on the IP solutions in an end-to-end manner. We analyze the feasibility of the IP procedure to determine the possible abstention rate for different levels of unfairness tolerance and accuracy constraint for achieving no harm. To the best of our knowledge, this work is the first to identify the theoretical relationships between the constraint parameters and the required abstention rate. Our theoretical results are important since a high abstention rate is often infeasible in practice due to a lack of human resources. Our framework outperforms existing methods in terms of fairness disparity without sacrificing accuracy at similar abstention rates. Tongxin Yin, Jean-Francois Ton, Ruocheng Guo, Yuanshun Yao, Mingyan Liu, Yang Liu 0018 |
ICLR | 4 |
| 2024 | Fairness without Harm: An Influence-Guided Active Sampling ApproachabstractThe pursuit of fairness in machine learning (ML), ensuring that the models do not exhibit biases toward protected demographic groups, typically results in a compromise scenario. This compromise can be explained by a Pareto frontier where given certain resources (e.g., data), reducing the fairness violations often comes at the cost of lowering the model accuracy.
In this work, we aim to train models that mitigate group fairness disparity without causing harm to model accuracy.
Intuitively, acquiring more data is a natural and promising approach to achieve this goal by reaching a better Pareto frontier of the fairness-accuracy tradeoff. The current data acquisition methods, such as fair active learning approaches, typically require annotating sensitive attributes. However, these sensitive attribute annotations should be protected due to privacy and safety concerns. In this paper, we propose a tractable active data sampling algorithm that does not rely on training group annotations, instead only requiring group annotations on a small validation set. Specifically, the algorithm first scores each new example by its influence on fairness and accuracy evaluated on the validation dataset, and then selects a certain number of examples for training.
We theoretically analyze how acquiring more data can improve fairness without causing harm, and validate the possibility of our sampling approach in the context of risk disparity. We also provide the upper bound of generalization error and risk disparity as well as the corresponding connections.
Extensive experiments on real-world data demonstrate the effectiveness of our proposed algorithm. Our code is available at [github.com/UCSC-REAL/FairnessWithoutHarm](https://github.com/UCSC-REAL/FairnessWithoutHarm). Jinlong Pang, Zhaowei Zhu, Yuanshun Yao, Chen Qian 0001, Yang Liu 0018 |
NeurIPS | 4 |
| 2024 | Large Language Model UnlearningabstractWe study how to perform unlearning, i.e. forgetting undesirable (mis)behaviors, on large language models (LLMs). We show at least three scenarios of aligning LLMs with human preferences can benefit from unlearning: (1) removing harmful responses, (2) erasing copyright-protected content as requested, and (3) reducing hallucinations. Unlearning, as an alignment technique, has three advantages. (1) It only requires negative (e.g. harmful) examples, which are much easier and cheaper to collect (e.g. via red teaming or user reporting) than positive (e.g. helpful and often human-written) examples required in the standard alignment process. (2) It is computationally efficient. (3) It is especially effective when we know which training samples cause the misbehavior. To the best of our knowledge, our work is among the first to explore LLM unlearning. We are also among the first to formulate the settings, goals, and evaluations in LLM unlearning. Despite only having negative samples, our ablation study shows that unlearning can still achieve better alignment performance than RLHF with just 2% of its computational time. Yuanshun Yao |
NeurIPS | 1 |
| 2023 | DPAUC: Differentially Private AUC Computation in Federated LearningabstractFederated learning (FL) has gained significant attention recently as a privacy-enhancing tool to jointly train a machine learning model by multiple participants. The prior work on FL has mostly studied how to protect label privacy during model training. However, model evaluation in FL might also lead to the potential leakage of private label information. In this work, we propose an evaluation algorithm that can accurately compute the widely used AUC (area under the curve) metric when using the label differential privacy (DP) in FL. Through extensive experiments, we show our algorithms can compute accurate AUCs compared to the ground truth. The code is available at https://github.com/bytedance/fedlearner/tree/master/example/privacy/DPAUC Jiankai Sun, Xin Yang 0017, Yuanshun Yao, Junyuan Xie, Chong Wang 0002 |
AAAI | 3 |
| 2023 | Weak Proxies are Sufficient and Preferable for Fairness with Missing Sensitive AttributesabstractEvaluating fairness can be challenging in practice because the sensitive attributes of data are often inaccessible due to privacy constraints. The go-to approach that the industry frequently adopts is using off-the-shelf proxy models to predict the missing sensitive attributes, e.g. Meta (Alao et al., 2021) and Twitter (Belli et al., 2022). Despite its popularity, there are three important questions unanswered: (1) Is directly using proxies efficacious in measuring fairness? (2) If not, is it possible to accurately evaluate fairness using proxies only? (3) Given the ethical controversy over infer-ring user private information, is it possible to only use weak (i.e. inaccurate) proxies in order to protect privacy? Our theoretical analyses show that directly using proxy models can give a false sense of (un)fairness. Second, we develop an algorithm that is able to measure fairness (provably) accurately with only three properly identified proxies. Third, we show that our algorithm allows the use of only weak proxies (e.g. with only 68.85% accuracy on COMPAS), adding an extra layer of protection on user privacy. Experiments validate our theoretical analyses and show our algorithm can effectively measure and mitigate bias. Our results imply a set of practical guidelines for prac-titioners on how to use proxies properly. Code is available at https://github.com/UCSC-REAL/fair-eval. Zhaowei Zhu, Yuanshun Yao, Jiankai Sun, Hang Li 0001, Yang Liu 0018 |
ICML | 2 |
| 2023 | "My face, my rules": Enabling Personalized Protection Against Unacceptable Face EditingabstractToday, face editing is widely used to refine/alter photos in both professional and recreational settings. Yet it is also used to modify (and repost) existing online photos for cyberbullying. Our work considers an important open question: 'How can we support the collaborative use of face editing on social platforms while protecting against unacceptable edits and reposts by others?' This is challenging because, as our user study shows, users vary widely in their definition of what edits are (un)acceptable. Any global filter policy deployed by social platforms is unlikely to address the needs of all users, but hinders social interactions enabled by photo editing. Instead, we argue that face edit protection policies should be implemented by social platforms based on individual user preferences. When posting an original photo online, a user can choose to specify the types of face edits (dis)allowed on the photo. Social platforms use these per-photo edit policies to moderate future photo uploads, i.e., edited photos containing modifications that violate the original photo's policy are either blocked or shelved for user approval. Realizing this personalized protection, however, faces two immediate challenges: (1) how to accurately recognize specific modifications, if any, contained in a photo; and (2) how to associate an edited photo with its original photo (and thus the edit policy). We show that these challenges can be addressed by combining highly efficient hashing based image search and scalable semantic image comparison, and build a prototype protector (Alethia) covering nine edit types. Evaluations using IRB-approved user studies and data-driven experiments (on 839K face photos) show that Alethia accurately recognizes edited photos that violate user policies and induces a feeling of protection to study participants. This demonstrates the initial feasibility of personalized face edit protection. We also discuss current limitations and future directions to push the concept forward. Zhujun Xiao, Jenna Cryan, Yuanshun Yao, Yi Hong Gordon Cheo, Yuanchao Shu, Stefan Saroiu, Ben Y. Zhao, Haitao Zheng 0001 |
Proc. Priv. Enhancing Technol. | 3 |
| 2022 | Differentially private multi-party data release for linear regressionabstractDifferentially Private (DP) data release is a promising technique to disseminate data without compromising the privacy of data subjects. However the majority of prior work has focused on scenarios where a single party owns all the data. In this paper we focus on the multi-party setting, where different stakeholders own disjoint sets of attributes belonging to the same group of data subjects. Within the context of linear regression that allow all parties to train models on the complete data without the ability to infer private attributes or identities of individuals, we start with directly applying Gaussian mechanism and show it has the small eigenvalue problem. We further propose our novel method and prove it asymptotically converges to the optimal (non-private) solutions with increasing dataset size. We substantiate the theoretical results through experiments on both artificial and real-world datasets. Ruihan Wu, Xin Yang 0017, Yuanshun Yao, Jiankai Sun, Kilian Q. Weinberger, Chong Wang 0002 |
UAI | 3 |
| 2021 | Backdoor Attacks Against Deep Learning Systems in the Physical WorldabstractBackdoor attacks embed hidden malicious behaviors into deep learning models, which only activate and cause misclassifications on model inputs containing a specific "trigger." Existing works on backdoor attacks and defenses, however, mostly focus on digital attacks that apply digitally generated patterns as triggers. A critical question remains unanswered: "can backdoor attacks succeed using physical objects as triggers, thus making them a credible threat against deep learning systems in the real world?"We conduct a detailed empirical study to explore this question for facial recognition, a critical deep learning task. Using 7 physical objects as triggers, we collect a custom dataset of 3205 images of 10 volunteers and use it to study the feasibility of "physical" backdoor attacks under a variety of real-world conditions. Our study reveals two key findings. First, physical backdoor attacks can be highly successful if they are carefully configured to overcome the constraints imposed by physical objects. In particular, the placement of successful triggers is largely constrained by the target model’s dependence on key facial features. Second, four of today’s state-of-the-art defenses against (digital) backdoors are ineffective against physical backdoors, because the use of physical objects breaks core assumptions used to construct these defenses.Our study confirms that (physical) backdoor attacks are not a hypothetical phenomenon but rather pose a serious real-world threat to critical classification tasks. We need new and more robust defenses against backdoors in the physical world. Emily Wenger, Josephine Passananti, Arjun Nitin Bhagoji, Yuanshun Yao, Haitao Zheng 0001, Ben Y. Zhao |
CVPR | 4 |
| 2019 | Latent Backdoor Attacks on Deep Neural NetworksabstractRecent work proposed the concept of backdoor attacks on deep neural networks (DNNs), where misclassification rules are hidden inside normal models, only to be triggered by very specific inputs. However, these "traditional" backdoors assume a context where users train their own models from scratch, which rarely occurs in practice. Instead, users typically customize "Teacher" models already pretrained by providers like Google, through a process called transfer learning. This customization process introduces significant changes to models and disrupts hidden backdoors, greatly reducing the actual impact of backdoors in practice. In this paper, we describe latent backdoors, a more powerful and stealthy variant of backdoor attacks that functions under transfer learning. Latent backdoors are incomplete backdoors embedded into a "Teacher" model, and automatically inherited by multiple "Student" models through transfer learning. If any Student models include the label targeted by the backdoor, then its customization process completes the backdoor and makes it active. We show that latent backdoors can be quite effective in a variety of application contexts, and validate its practicality through real-world attacks against traffic sign recognition, iris identification of volunteers, and facial recognition of public figures (politicians). Finally, we evaluate 4 potential defenses, and find that only one is effective in disrupting latent backdoors, but might incur a cost in classification accuracy as tradeoff. Yuanshun Yao, Huiying Li 0001, Haitao Zheng 0001, Ben Y. Zhao |
CCS | 1 |
| 2019 | Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksabstractLack of transparency in deep neural networks (DNNs) make them susceptible to backdoor attacks, where hidden associations or triggers override normal classification to produce unexpected results. For example, a model with a backdoor always identifies a face as Bill Gates if a specific symbol is present in the input. Backdoors can stay hidden indefinitely until activated by an input, and present a serious security risk to many security or safety related applications, e.g. biometric authentication systems or self-driving cars. We present the first robust and generalizable detection and mitigation system for DNN backdoor attacks. Our techniques identify backdoors and reconstruct possible triggers. We identify multiple mitigation techniques via input filters, neuron pruning and unlearning. We demonstrate their efficacy via extensive experiments on a variety of DNNs, against two types of backdoor injection methods identified by prior work. Our techniques also prove robust against a number of variants of the backdoor attack. Bolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li 0001, Bimal Viswanath, Haitao Zheng 0001, Ben Y. Zhao |
IEEE Symposium on Security and Privacy | 2 |
| 2018 | With Great Training Comes Great Vulnerability: Practical Attacks against Transfer Learning
Bolun Wang, Yuanshun Yao, Bimal Viswanath, Haitao Zheng 0001, Ben Y. Zhao |
USENIX Security Symposium | 2 |
| 2017 | Automated Crowdturfing Attacks and Defenses in Online Review SystemsabstractMalicious crowdsourcing forums are gaining traction as sources of spreading misinformation online, but are limited by the costs of hiring and managing human workers. In this paper, we identify a new class of attacks that leverage deep learning language models (Recurrent Neural Networks or RNNs) to automate the generation of fake online reviews for products and services. Not only are these attacks cheap and therefore more scalable, but they can control rate of content output to eliminate the signature burstiness that makes crowdsourced campaigns easy to detect. Yuanshun Yao, Bimal Viswanath, Jenna Cryan, Haitao Zheng 0001, Ben Y. Zhao |
CCS | 1 |
| 2017 | Complexity vs. performance: empirical analysis of machine learning as a serviceabstractMachine learning classifiers are basic research tools used in numerous types of network analysis and modeling. To reduce the need for domain expertise and costs of running local ML classifiers, network researchers can instead rely on centralized Machine Learning as a Service (MLaaS) platforms. Yuanshun Yao, Zhujun Xiao, Bolun Wang, Bimal Viswanath, Haitao Zheng 0001, Ben Y. Zhao |
Internet Measurement Conference | 1 |
| 2017 | Object Recognition and Navigation using a Single Networking DeviceabstractTomorrow's autonomous mobile devices need accurate, robust and real-time sensing of their operating environment. Today's solutions fall short. Vision or acoustic-based techniques are vulnerable against challenging lighting conditions or background noise, while more robust laser or RF solutions require either bulky expensive hardware or tight coordination between multiple devices. This paper describes the design, implementation and evaluation of Ulysses, a practical environmental imaging system using colocated 60GHz radios on a single mobile device. Unlike alternatives that require specialized hardware, Ulysses reuses low-cost commodity networking chipsets available today. Ulysses' new imaging approach leverages RF beamforming, operates on specular (direct) reflection, and integrates the device's movement trajectory with sensing. Ulysses also includes a navigation component that uses the same 60GHz radios to compute "safety regions" where devices can move freely without collision, and to compute optimal paths for imaging within safety regions. Using our implementation of a small robotic car prototype, our experimental results show that Ulysses images objects meters away with cm-level precision, and provides accurate estimates of objects' surface materials. Yanzi Zhu, Yuanshun Yao, Ben Y. Zhao, Haitao Zheng 0001 |
MobiSys | 2 |
| 2017 | Identifying Value in Crowdsourced Wireless Signal MeasurementsabstractWhile crowdsourcing is an attractive approach to collect large-scale wireless measurements, understanding the quality and variance of the resulting data is difficult. Our work analyzes the quality of crowdsourced cellular signal measurements in the context of basestation localization, using large international public datasets (419M signal measurements and 1M cells) and corresponding ground truth values. Performing localization using raw received signal strength (RSS) data produces poor results and very high variance. Applying supervised learning improves results moderately, but variance remains high. Instead, we propose feature clustering, a novel application of unsupervised learning to detect hidden correlation between measurement instances, their features, and localization accuracy. Our results identify RSS standard deviation and RSS-weighted dispersion mean as key features that correlate with highly predictive measurement samples for both sparse and dense measurements respectively. Finally, we show how optimizing crowdsourcing measurements for these two features dramatically improves localization accuracy and reduces variance. Zhijing Li 0001, Ana Nika, Xinyi Zhang 0003, Yanzi Zhu, Yuanshun Yao, Ben Y. Zhao, Haitao Zheng 0001 |
WWW | 5 |
| 2016 | A general framework to increase the robustness of model-based change point detection algorithms to outliers and noiseabstractThe autonomous identification of time-steps where the behavior of a time-series significantly deviates from a predefined model, or time-series change point detection, is an active field of research with notable applications in finance, health, and advertising. One family of time-series change detection algorithms, referred to as “model-based methods”, although useful for many applications, performs poor when the data are noisy and have outliers. We introduce a new framework that enables existing model-based methods to be more robust to these data challenges. We demonstrate the effectiveness of our approach on remote sensing and mobile health data. Our method introduces two new concepts: (i) a random sampling procedure allows us to overcome outliers, and (ii) a matrix-based representation of anomaly scores provides a flexible and intuitive way to identify multiple types of changes and test their significance. We show that our method performs better than several baseline methods, including application-specific algorithms, and provide all data and open-source code. Xi Chen 0120, Yuanshun Yao, Sichao Shi, Snigdhansu Chatterjee, Vipin Kumar 0001, James H. Faghmous |
SDM | 2 |