VLDB 2026 Research / reviewers in the wild / expert
Haitao Xu 0002
dblp:41/10114-2
· DBLP profile ↗
32ranked-venue papers
8as first author
22since 2021 · last 2026
0000-0002-0353-3879ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 15 · 4 first-author · 10 since 2021Computer networks · 7 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LiBrain: LLM-Powered Li-ion Battery Diagnostics with Time-Series-Aware Retrieval-Augmented Framework for E-bikes
Zhao Li 0007, Zixin Lin, Donghui Ding, Yichen Zhong, Haitao Xu 0002, Peng Cai 0001 |
AAAI | 6 |
| 2026 | The Privacy Paradox of LLMs: User Perceptions and the Reality of PII LeakageabstractLarge language models (LLMs) are increasingly deployed, yet they introduce significant privacy risks by disclosing personally identifiable information (PII) during interactions. Although prior work has demonstrated the feasibility of extracting PII from LLMs, no comprehensive study has evaluated the actual extent of PII leakage across mainstream LLMs or investigated user perceptions, literacy, and behavioral responses to these risks. To address these gaps, we conduct a large-scale evaluation of PII leakage in popular LLMs, demonstrating that attackers can extract email addresses and phone numbers with high success rates. Through a mixed-methods study involving 20 interviews and 204 survey participants, we identify significant discrepancies between user concerns and behavior: despite strong concerns about PII leakage and limited understanding of training data provenance, users continue to use LLMs due to perceived utility, often exhibiting privacy cynicism. Based on these findings, we propose design implications for enhancing the privacy-utility balance in future LLM deployments. Haitao Xu 0002, Shu Meng, Shuai Hao 0001, Chuan Yue, Zhao Li 0007 |
CHI | 2 |
| 2026 | ENDeliver: An Energy-Aware Framework of Large Spatio-Temporal Model for E-Bike Delivery Route Planning
Yuduo Shi, Zhao Li 0007, Wenrui Ma, Haitao Xu 0002 |
DASFAA (6) | 4 |
| 2026 | LLM-Empowered Discovery of Windows APIs Exploitable for Persistent Storage in Fileless Attacks
Shu Meng, Haitao Xu 0002, Shuai Hao 0001, Yixin Jiang |
DSN | 3 |
| 2026 | CHAMELEOSCAN: Demystifying and Detecting iOS Chameleon Apps via LLM-Powered UI Exploration
Haitao Xu 0002, Yanchen Lu, Mengxia Ren, Shuai Hao 0001, Chuan Yue, Zhao Li 0007, Fan Zhang 0010, Yixin Jiang |
NDSS | 3 |
| 2026 | Incorporating Gradients to Rules: Toward Online, Adaptive Provenance-Based Intrusion DetectionabstractAs cyber-attacks become increasingly sophisticated and stealthy, accurately distinguishing between benign behavior and malicious intrusions has become both more critical and more challenging. Provenance-based intrusion detection systems (PIDS) show strong potential for detecting malicious activities through fine-grained causality analysis, which has gained significant attention from both industry and academia. Among the various PIDS approaches, rule-based systems are particularly favored for their low overhead, real-time detection capability, and interpretability. However, these systems face challenges in reducing false positive rates, primarily due to the lack of fine-tuned rules and specific environments. In this paper, we introduce CAPTAIN+, a rule-based PIDS that autonomously adapts to diverse environments online. Specifically, we propose three adaptive parameters to adjust the detection configuration for nodes, edges, and alarm generation thresholds. Initially, we build a differentiable tag propagation framework and utilize the gradient descent algorithm to optimize these adaptive parameters based on the training data. In this extended version, we integrate an online learning module into the detection stage to dynamically optimize adaptive parameters based on real-time feedback from the detection process. We evaluate CAPTAIN+ based on data from DARPA TC, OpTC datasets, and PKU ASAL datasets. The results demonstrate that CAPTAIN+ offers superior detection accuracy, lower detection latency, reduced runtime overhead, long-term resilience against concept drift, and more interpretable detection results compared to state-of-the-art PIDS. Zhenyuan Li, Lingzhi Wang 0002, Zhengkai Wang, Xiangmin Shen, Haitao Xu 0002, Yan Chen 0004, Shouling Ji |
IEEE Trans. Dependable Secur. Comput. | 5 |
| 2025 | TacDroid: Detection of Illicit Apps Through Hybrid Analysis of UI-Based Transition GraphsabstractIllicit apps have emerged as a thriving underground industry, driven by their substantial profitability. These apps either offer users restricted services (e.g., porn and gambling) or engage in fraudulent activities like scams. Despite the widespread presence of illicit apps, scant attention has been directed towards this issue, with several existing detection methods predominantly relying on static analysis alone. However, given the burgeoning trend wherein an increasing number of mobile apps achieve their core functionality through dynamic resource loading, depending solely on static analysis proves inadequate. To address this challenge, in this paper, we introduce Tac-droid,a novel approach that integrates dynamic analysis for dynamic content retrieval with static analysis to mitigate the limitations inherent in both methods, i.e., the low coverage of dynamic analysis and the low accuracy of static analysis. Specifically, Tacdroid conducts both dynamic and static analyses on an Android app to construct dynamic and static User Interface Transition Graphs (UTGs), respectively. These two UTGs are then correlated to create an intermediate UTG. Subsequently, Tacdroid embeds graph structure and utilizes an enhanced Graph Autoencoder (GAE) model to predict transitions between nodes. Through link prediction, Tacdroid effectively eliminates false positive transition edges stemming from misjudgments in static analysis and supplements false negative transition edges overlooked in the intermediate UTG, thereby generating a comprehensive and accurate UTG. Finally, Tacdroid determines the legitimacy of an app and identifies its category based on the app's UTG. Our evaluation results highlight the outstanding accuracy of Tacdroid in detecting illicit apps. It significantly surpasses the state-of-the-art work, achieving an F1-score of 96.73%. This work represents a notable advancement in the identification and categorization of illicit apps. Yanchen Lu, Zehua He, Haitao Xu 0002, Zhao Li 0007, Shuai Hao 0001, Liu Wang 0002, Haoyu Wang 0001, Kui Ren 0001 |
ICSE | 4 |
| 2025 | Understanding PII Leakage in Large Language Models: A Systematic SurveyabstractLarge Language Models (LLMs) have demonstrated exceptional success across a variety of tasks, particularly in natural language processing, leading to their growing integration into numerous facets of daily life. However, this widespread deployment has raised substantial privacy concerns, especially regarding personally identifiable information (PII), which can be directly associated with specific individuals. The leakage of such information presents significant real-world privacy threats. In this paper, we conduct a systematic investigation into existing research on PII leakage in LLMs, encompassing commonly utilized PII datasets, evaluation metrics, and current studies on both PII leakage attacks and defensive strategies. Finally, we identify unresolved challenges in the current research landscape and suggest future research directions. Zhao Li 0007, Shu Meng, Mengxia Ren, Haitao Xu 0002, Shuai Hao 0001, Chuan Yue, Fan Zhang 0010 |
IJCAI | 5 |
| 2025 | Understanding the Business of Online Affiliate Marketing: An Empirical StudyabstractAffiliate marketing is a revenue-sharing marketing scheme by which an affiliate, such as a blogger or YouTuber, garners commissions for promoting a merchant's goods or services, thereby aiming to foster a mutually beneficial relationship between affiliates and merchants. Despite being a multi-billion-dollar global industry, affiliate marketing remains inadequately explored, and the research community lacks a comprehensive understanding of its intricate ecosystem. In this paper, we present the first comprehensive empirical study of the affiliate marketing ecosystem. We conduct thorough measurements to assess the prevalence of affiliate marketing, estimate the market size, and elucidate the characteristics of affiliates, merchants, and intermediary affiliate networks. Over a continuous span of 13 months, we monitored four of the most prominent affiliate aggregation platforms, yielding a substantial dataset. We observed 467,219 unique offers - tasks to be undertaken by affiliates - involving 37,109 merchants and 556 affiliate networks across the four platforms. Notably, these offers would cost the merchants more than 19 million USD for the completion of all the actions pre-defined in these offers, such as signing up or making a transaction. Additionally, we compiled a large-scale dataset comprising 124,462 affiliate links, enabling us to conduct a comprehensive investigation. Finally, we propose machine learning models incorporating the characteristics of affiliate links to detect real-world affiliate marketing campaigns. Haitao Xu 0002, Kaleem Ullah Qasim, Shuai Hao 0001, Wenrui Ma, Zhenyuan Li, Fan Zhang 0010, Zhao Li 0007 |
INFOCOM | 1 |
| 2025 | Spatiotemporal Cross-Domain Integrated Insights: Mitigating Fraudulent Activities on Ethereum
Yuduo Shi, Zhao Li 0007, Haitao Xu 0002, Wenrui Ma, Ji Zhang 0001 |
SecureComm (5) | 3 |
| 2025 | Effective PII Extraction from LLMs through Augmented Few-Shot Learning
Shu Meng, Haitao Xu 0002, Shuai Hao 0001, Chuan Yue, Wenrui Ma, Fan Zhang 0010, Zhao Li 0007 |
USENIX Security Symposium | 3 |
| 2024 | DDoSMiner: An Automated Framework for DDoS Attack Characterization and Vulnerability Mining
Xi Ling, Jiongchi Yu, Ziming Zhao 0008, Haitao Xu 0002, Binbin Chen 0001, Fan Zhang 0010 |
ACNS (2) | 5 |
| 2024 | Informative Sample Labeling with Conditional Variational Deep Embedding for Active Learning
Zhao Li 0007, Qinxue Meng, Haitao Xu 0002, Yangbohan Jiao, Buqing Cao |
ADMA (1) | 3 |
| 2024 | Android Malware Family Labeling: Perspectives from the IndustryabstractLabeling and classifying Android malware is important for identifying new threats, triaging security incidents, and demystifying evasion techniques. To automate the malware classification pipeline, state-of-the-art tools such as AVClass and Euphony unify raw labels from commercial antivirus vendors (i.e., VirusTotal) to produce family labels. These tools are widely used for automatic malware classification in both academic research and industry practice. However, they face significant limitations in real-world industrial scenarios with numerous and dynamically changing samples. For example, our industrial practices revealed that VirusTotal's results change over time, leading to temporal inconsistencies in family labeling results that rely on label unification, which can severely impact a company's security posture. Despite this, such issues and challenges remain understudied. In this paper, we present the first systematic measurement study of existing automatic Android malware family labeling systems from various aspects, including label dynamics, consistency, reliability, and etc. Based on a large-scale dataset, we validate that the labeling results of these systems do evolve with time, and such evolution can introduce bias into many previous studies on performance assessments. We also reveal substantial divergence in labeling decisions across different systems when given the same input. Besides, we identify a disclosure priority among families in these systems' labeling processes, which could threaten the industry by allowing malicious actors to exploit these discrepancies. Our findings could benefit both researchers and industry practitioners for further refinement of automatic malware family labeling systems, contributing to their practical applications. Liu Wang 0002, Haoyu Wang 0001, Tao Zhang 0001, Haitao Xu 0002, Guozhu Meng, Peiming Gao, Yi Wang 0013 |
ASE | 4 |
| 2024 | TransURL: Improving malicious URL detection with multi-layer Transformer encoding and multi-scale pyramid features
Zhenhao Guo, Haitao Xu 0002, Zhan Qin, Wenrui Ma, Fan Zhang 0010 |
Comput. Networks | 4 |
| 2024 | CMD: Co-Analyzed IoT Malware Detection and Forensics via Network and Hardware DomainsabstractWith the widespread use of Internet of Things (IoT) devices, malware detection has become a hot spot for both academic and industrial communities. Existing approaches can be roughly categorized into network-side and host-side. However, existing network-side methods are difficult to capture contextual semantics from cross-source traffic, and previous host-side methods could be adversary-perceived and expose risks for tampering. More importantly, a single perspective cannot comprehensively track the multi-stage lifecycle of IoT malware. In this paper, we present${\sf CMD}$, a co-analyzed IoT malware detection and forensics system by combining hardware and network domains. For the network part,${\sf CMD}$proposes a tailored capsule neural network to capture the contextual semantics from cross-source traffic. For the hardware part,${\sf CMD}$designs an entire file operation recovery process in a side-channel manner by leveraging the Serial Peripheral Interface (SPI) signals from on-chip traces. These traffic provenance and operating logs information could benefit the anti-virus countermeasures for security practitioners. By practical evaluation, we demonstrate that${\sf CMD}$realizes outstanding detection effects (e.g.,$\sim$99.88% F1-score) compared with seven state-of-the-art methods, and recovers 96.88%$\sim$99.75% operation commands even if against adaptive adversaries (that could kill processes or tamper with operation log files). A by-product benefit of such an external monitor is${\sf CMD}$introduces zero latency on the IoT device, and incurs negligible IoT CPU utilization. Also, since SPI focuses on file operations, the proposed hardware trace forensics does not have the data explosion problem like previous work,e.g.,recovered logs of${\sf CMD}$only take up limited extra space overhead (e.g.,$\sim$0.2 MB per malware). Furthermore, we provide the model interpretability for the capsule network and develop a case study (Hajime) of the operation logs recovery. Ziming Zhao 0008, Zhaoxuan Li, Jiongchi Yu, Fan Zhang 0010, Xiaofei Xie, Haitao Xu 0002, Binbin Chen 0001 |
IEEE Trans. Mob. Comput. | 6 |
| 2023 | A Large-Scale Pretrained Deep Model for Phishing URL DetectionabstractPhishing attacks have always been a security issue that has attracted great attention in the cyber security community. Recently, the famous pre-trained models is being used as an anti-phishing solution. However, existing studies either simply transfer models pre-trained on text to phishing detection task, or pre-train models using only extremely small phishing samples. In this paper, we propose PhishBERT, a veritable pretrained deep transformer network model for phishing URL detection. Using a tailor pre-training objective, PhishBERT obtained a general understanding of various URLs by being pretrained on a corpus of more than 3 billion unlabeled URL data. It is then transferred to the detection task of benign and malicious URL data, with supervised fine-tuning using adversarial methods. Extensive and rigorous benchmark studies verify that PhishBERT is significantly superior to the current state-of-the-art methods in terms of efficiency, robustness and accuracy on the task of phishing website detection. Weifan Zhu, Haitao Xu 0002, Zhan Qin, Kui Ren 0001, Wenrui Ma |
ICASSP | 3 |
| 2023 | Investigating Fraud and Misconduct in Legitimate Internet Economy based on Customer ComplaintsabstractDifferent forms of cybercrimes, ranging from email spam and click fraud to the most sophisticated underground economy, have been studied extensively. Fraud and misconduct in legitimate Internet economy, however, have not received sufficient attention, even though potential damages may not be as severe as regular cybercrimes. In this paper, we have performed the first in-depth empirical investigation of fraud and misconduct in legitimate Internet businesses by collecting 6.6 million customer complaints, which were filed over 45 months by 2.7 million customers against 117 thousand merchants on one of the largest customer complaint platforms in the world, along with more than 13 million customer-uploaded images as photographic evidence. We characterized the complaints in terms of merchants, main issues, desired remedy, amount of money involved, customer ratings, and etc. Most importantly, we were able to uncover various fraud and misconduct in different Internet businesses, some of which should have caught law enforcement's attention. The sum of the amount of money involved in these complaints is 5.4 billion US dollars. We also investigated the privacy inference of the images, and found that the image content could disclose an unexpected amount of customers' personal information. Wenrui Ma, Ying Cong, Haitao Xu 0002, Fan Zhang 0010, Zhao Li 0007, Siqi Ren |
TrustCom | 3 |
| 2022 | RATScope: Recording and Reconstructing Missing RAT Semantic Behaviors for Forensic Analysis on WindowsabstractRemote Access Trojan (RAT) attacks have become an extensively prevailing and serious threat to enterprise security. A forensic system targeting RAT attacks is needed to record and reconstruct fine-grained semantic behaviors of RATs. However, existing forensic systems suffer from various issues such as intrusive instrumentation, nontrivial recording overhead, and RAT behavior blindness. In this article, we first conduct a large-scale study of a representative set of real-world RAT families active from 1999 to 2016. This is the first study to understand the landscape of RATs in the literature. Based on the study, we then proposeRATScope, an instrumentation-free RAT forensic system targeting Windows platform. Specifically,RATScopeoffers an audit logging module to efficiently record system logs by leveraging Event Tracing for Windows (ETW), and provides a novel program behavior modeling technique to reconstruct semantic behaviors of RATs accurately. We implement a prototype ofRATScopeand evaluate the recording overhead and the behavior identification accuracy. The results show that the audit logging module only incurs 3.7 percent runtime overhead on average. Our system can achieve around 90 percent true positive rate in the cross-family experiment, around 80 percent true positive rate in the two-year spanning temporal experiment, and nearzerofalse positive rate. Runqing Yang, Xutong Chen, Haitao Xu 0002, Yueqiang Cheng, Chun-lin Xiong, Linqi Ruan, Mohammad Kavousi, Zhenyuan Li, Liheng Xu, Yan Chen 0004 |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2022 | snWF: Website Fingerprinting Attack by Ensembling the Snapshot of Deep LearningabstractThe website fingerprinting (WF) attack enables a local eavesdropper to identify which website a client is visiting under encrypted network connections. By leveraging deep neural networks, the state-of-the-art WF attacks achieve high accuracy in classic experimental scenes. However, due to the high variance of neural networks, those attacks are sensitive to the specific information in the data and would result in less-than-ideal performance on data outside the training set or on data that is impacted by the effect of concept drift. In this paper, we present snWF, a novelWF attack, which leverages an out-of-the ordinary ensemble to reduce the variance of neural networks and improve the robustness of the attack. In a large open-world setting with 400,000 websites, snWF manages to determine whether a user is visiting a monitored website, with a true positive rate of 98.1% and a false positive rate of 5.7%. We also evaluated snWF in a more realistic attack scenario, termed aswide-world, to examine whether snWF can correctly classify websites that even an adversary has not seen before, and we found that snWF achieves a higher classification accuracy than the state-of-the-art attacks in this new setting. In addition, in the face of concept drift, snWF is found to be more resilient than any other attacks. Moreover, we are the first to reveal that under concept drift WF attacks suffer more severe performance degradation in aopen-worldsetting than in aclosed-worldsetting. Haitao Xu 0002, Zhenhao Guo, Zhan Qin, Kui Ren 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2021 | A Study of the Partnership Between Advertisers and Publishers
Wenrui Ma, Haitao Xu 0002 |
PAM | 2 |
| 2021 | MAdLens: Investigating Into Android In-App Ad Practice at API GranularityabstractIn-app advertising has served as the major revenue source for millions of app developers in the mobile Internet ecosystem. Ad networks play an important role in app monetization by providing third-party libraries for developers to choose and embed into their apps. Various ad mediations help developers manage all of the ad libraries used in apps to show the best available ad among received ads from different ad network servers. However, developers lack guidelines on how to choose from hundreds of ad networks or ad mediations and various ad features to maximize their revenues without hurting the user experience of their apps. Our work aims to provide app developers guidelines on the selection of ad networks, ad mediations, and ad placement by observing current common practices. To this end, we investigate 838 unique APIs from 207 ad networks which are extracted from 277,616 Android apps, develop a methodology of ad type classification based on UI interaction and behavior, and perform a large scale measurement study of in-app ads with static analysis techniques at the API granularity. We found that developers have more choices about ad networks than several years before. Most developers are conservative about ad placement and about 77 percent of the apps contain at most one ad library. Besides, the likeliness of an app containing ads depends on the app category to which it belongs. Furthermore, we propose a terminology and classify mobile ads into five ad types: Embedded, Popup, Notification, Offerwall, and Floating. Also, our research shows that it is a better solution for developers to integrate ad libraries with ad mediation feature in their apps because it may avoid bad ratings and improve user experience. And in our findings, more than 95 percent of embedded, popup, notification, and offer ads locate in the zero activity (main activity), the first activity and the second activity of Android apps. More interestingly, developers tend to put high aggressive ads on activities which need deeper user interaction. Our research is the first to reveal the preference of both developers and users for ad networks, ad mediation feature and ad types. Ling Jin 0005, Boyuan He, Guangyao Weng, Haitao Xu 0002, Yan Chen 0004, Guanyu Guo |
IEEE Trans. Mob. Comput. | 4 |
| 2020 | UIScope: Accurate, Instrumentation-free, and Visible Attack Investigation for GUI Applications
Runqing Yang, Shiqing Ma, Haitao Xu 0002, Xiangyu Zhang 0001, Yan Chen 0004 |
NDSS | 3 |
| 2020 | RiskCog: Unobtrusive Real-Time User Authentication on Mobile Devices in the WildabstractRecent hardware advances have led to the development and consumerization of mobile devices, which mainly include smartphones and various wearable devices. To protect the privacy of users, various user authentication mechanisms have been proposed. In particular, biometrics has been widely used for multi-factor authentication. However, biometrics-based authentication mechanisms usually require costly sensors deployed on devices, and rely on explicit user input and Internet connection for performing user authentication. In this article, we propose a system, called RISKCOG, which can authenticate the ownership of mobile devices unobtrusively and in a real-time manner by adopting a learning-based approach. Unlike previous studies on user authentication, for cross-platform deployment, maximum user privacy protection, and unobtrusive authentication, RISKCOG only relies on those widely available and privacy-insensitive motion sensors to capture the data related to the users' daily device usage. It requires no users' explicit input and has no requirement on the users' motion state or the device placement. RISKCOG is also usable in the environment without Internet access by performing offline user identity verification. We conduct comprehensive experiments on smartphones and smartwatches, which show that RISKCOG can authenticate device users rapidly and with high accuracy. Tiantian Zhu 0001, Zhengyang Qu, Haitao Xu 0002, Jingsi Zhang, Zhengyue Shao, Yan Chen 0004, Sandeep Prabhakar |
IEEE Trans. Mob. Comput. | 3 |
| 2018 | Detecting and Characterizing Web Bot Traffic in a Large E-commerce Marketplace
Haitao Xu 0002, Zhao Li 0007, Chen Chu, Yuanmi Chen, Yifan Yang 0001, Haifeng Lu, Haining Wang 0001, Angelos Stavrou |
ESORICS (2) | 1 |
| 2018 | An Investigation into Android In-App Ad Practice: Implications for App DevelopersabstractIn-app advertising has served as the major revenue source for millions of app developers in the mobile Internet ecosystem. Ad networks play an important role in app monetization by providing third-party libraries for developers to choose and embed into their apps. However, developers lack guidelines on how to choose from hundreds of ad networks and various ad features to maximize their revues without hurting the user experience of their apps. Our work aims to uncover the best practice and provide app developers guidelines on ad network selection and ad placement. To this end, we investigate 697 unique APIs from 164 ad networks which are extracted from 277,616 Android apps, develop a methodology of ad type classification based on UI interaction and behavior, and perform a large scale measurement study of in-app ads with static analysis techniques at the API granularity. We found that developers have more choices about ad networks than several years before. Most developers are conservative about ad placement and about 71% apps contain at most one ad library. In addition, the likeliness of an app containing ads depends on the app category to which it belongs. The app categories featuring young audience usually contain the most ad libraries maybe because of the ad-tolerance characteristic of young people. Furthermore, we propose a terminology and classify mobile ads into five ad types: Embedded, Popup, Notification, Offerwall, and Floating. We found that embedded and popup ad types are popular with apps in nearly all categories. Our results also suggest that developers should embed at most 6 ad libraries into an app, which otherwise would anger the app users. Also, a developer should use at most one ad network when her app is still at the initial stage and could start using more (2 or 3) ad networks when the app becomes popular. Our research is the first to reveal the preference of both developers and users for ad networks and ad types. Boyuan He, Haitao Xu 0002, Ling Jin 0005, Guanyu Guo, Yan Chen 0004, Guangyao Weng |
INFOCOM | 2 |
| 2018 | Privacy Risk Assessment on Email TrackingabstractToday's online marketing industry has widely employed email tracking techniques, such as embedding a tiny tracking pixel, to track email opens of potential customers and measure marketing effectiveness. However, email tracking could allow miscreants to collect metadata information associated with email reading without user awareness and then leverage the information for stealthy surveillance, which has raised serious privacy concerns. In this paper, we present an in-depth and comprehensive study on the privacy implications of email tracking. First, we develop an email tracking system and perform realworld tracking on hundreds of solicited crowdsourcing participants. We estimate the amount of privacy-sensitive information available from email reading, assess privacy risks of information leakage, and demonstrate how easy it is to launch a long-term targeted surveillance attack in real scenarios by simply sending an email with tracking capability. Second, we investigate the prevalence of email tracking through a large-scale measurement, which includes more than 44,000 email samples obtained over a period of seven years. Third, we conduct a user study to understand users' perception of privacy infringement caused by email tracking. Finally, we evaluate existing countermeasures against email tracking and propose guidelines for developing more comprehensive and fine-grained prevention solutions. Haitao Xu 0002, Shuai Hao 0001, Alparslan Sari, Haining Wang 0001 |
INFOCOM | 1 |
| 2018 | Internet Protocol Cameras with No Password Protection: An Empirical Investigation
Haitao Xu 0002, Fengyuan Xu, Bo Chen 0028 |
PAM | 1 |
| 2017 | An Empirical Investigation of Ecommerce-Reputation-Escalation-as-a-ServiceabstractIn online markets, a store’s reputation is closely tied to its profitability. Sellers’ desire to quickly achieve a high reputation has fueled a profitable underground business that operates as a specialized crowdsourcing marketplace and accumulates wealth by allowing online sellers to harness human laborers to conduct fake transactions to improve their stores’ reputations. We term such an underground market a seller-reputation-escalation (SRE) market . In this article, we investigate the impact of the SRE service on reputation escalation by performing in-depth measurements of the prevalence of the SRE service, the business model and market size of SRE markets, and the characteristics of sellers and offered laborers. To this end, we have infiltrated five SRE markets and studied their operations using daily data collection over a continuous period of 2 months. We identified more than 11,000 online sellers posting at least 219,165 fake-purchase tasks on the five SRE markets. These transactions earned at least $46,438 in revenue for the five SRE markets, and the total value of merchandise involved exceeded $3,452,530. Our study demonstrates that online sellers using the SRE service can increase their stores’ reputations at least 10 times faster than legitimate ones while about 25% of them were visibly penalized. Even worse, we found a much stealthier and more hazardous service that can, within a single day, boost a seller’s reputation by such a degree that would require a legitimate seller at least a year to accomplish. Armed with our analysis of the operational characteristics of the underground economy, we offer some insights into potential mitigation strategies. Finally, we revisit the SRE ecosystem 1 year later to evaluate the latest dynamism of the SRE markets, especially the statuses of the online stores once identified to launch fake-transaction campaigns on the SRE markets. We observe that the SRE markets are not as active as they were 1 year ago and about 17% of the involved online stores become inaccessible likely because they have been forcibly shut down by the corresponding E-commerce marketplace for conducting fake transactions. Haitao Xu 0002, Daiping Liu, Haining Wang 0001, Angelos Stavrou |
ACM Trans. Web | 1 |
| 2015 | Privacy Risk Assessment on Online Photos
Haitao Xu 0002, Haining Wang 0001, Angelos Stavrou |
RAID | 1 |
| 2015 | E-commerce Reputation Manipulation: The Emergence of Reputation-Escalation-as-a-ServiceabstractIn online markets, a store's reputation is closely tied to its profitability. Sellers' desire to quickly achieve high reputation has fueled a profitable underground business, which operates as a specialized crowdsourcing marketplace and accumulates wealth by allowing online sellers to harness human laborers to conduct fake transactions for improving their stores' reputations. We term such an underground market a seller-reputation-escalation (SRE) market. In this paper, we investigate the impact of the SRE service on reputation escalation by performing in-depth measurements of the prevalence of the SRE service, the business model and market size of SRE markets, and the characteristics of sellers and offered laborers. To this end, we have infiltrated five SRE markets and studied their operations using daily data collection over a continuous period of two months. We identified more than 11,000 online sellers posting at least 219,165 fake-purchase tasks on the five SRE markets. These transactions earned at least $46,438 in revenue for the five SRE markets, and the total value of merchandise involved exceeded $3,452,530. Our study demonstrates that online sellers using SRE service can increase their stores' reputations at least 10 times faster than legitimate ones while only 2.2% of them were detected and penalized. Even worse, we found a newly launched service that can, within a single day, boost a seller's reputation by such a degree that would require a legitimate seller at least a year to accomplish. Finally, armed with our analysis of the operational characteristics of the underground economy, we offer some insights into potential mitigation strategies. Haitao Xu 0002, Daiping Liu, Haining Wang 0001, Angelos Stavrou |
WWW | 1 |
| 2014 | Click Fraud Detection on the Advertiser Side
Haitao Xu 0002, Daiping Liu, Aaron Koehl, Haining Wang 0001, Angelos Stavrou |
ESORICS (2) | 1 |