VLDB 2026 Research / reviewers in the wild / expert
Yubao Zhang
dblp:65/8966
· DBLP profile ↗
14ranked-venue papers
6as first author
5since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 10 · 5 first-author · 3 since 2021Computer networks · 3 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | A multiview clustering framework for detecting deceptive reviewsabstractOnline reviews, which play a key role in the ecosystem of nowadays business, have been the primary source of consumer opinions. Due to their importance, professional review writing services are employed for paid reviews and even being exploited to conduct opinion spam. Posting deceptive reviews could mislead customers, yield significant benefits or losses to service vendors, and erode confidence in the entire online purchasing ecosystem. In this paper, we ferret out deceptive reviews originated from professional review writing services. We do so even when reviewers leverage a number of pseudonymous identities to avoid the detection. To unveil the pseudonymous identities associated with deceptive reviewers, we leverage the multiview clustering method. This enables us to characterize the writing style of reviewers (deceptive vs normal) and cluster the reviewers based on their writing style. Furthermore, we explore different neural network models to model the writing style of deceptive reviews. We select the best performing neural network to generate the representation of reviews. We validate the effectiveness of the multiview clustering framework using real-world Amazon review data under different experimental scenarios. Our results show that our approach outperforms previous research. We further demonstrate its superiority through a large-scale case study based on publicly available Amazon datasets. Yubao Zhang, Haining Wang 0001, Angelos Stavrou |
J. Comput. Secur. | 1 |
| 2023 | Dial "N" for NXDomain: The Scale, Origin, and Security Implications of DNS Queries to Non-Existent DomainsabstractNon-Existent Domain (NXDomain) is one type of the Domain Name System (DNS) error responses, indicating that the queried domain name does not exist and cannot be resolved. Unfortunately, little research has focused on understanding why and how NXDomain responses are generated, utilized, and exploited. In this paper, we conduct the first comprehensive and systematic study on NXDomain by investigating its scale, origin, and security implications. Utilizing a large-scale passive DNS database, we identify 146,363,745,785 NXDomains queried by DNS users between 2014 and 2022. Within these 146 billion NXDomains, 91 million of them hold historic WHOIS records, of which 5.3 million are identified as malicious domains including about 2.4 million blocklisted domains, 2.8 million DGA (Domain Generation Algorithms) based domains, and 90 thousand squatting domains targeting popular domains. To gain more insights into the usage patterns and security risks of NXDomains, we register 19 carefully selected NXDomains in the DNS database, each of which received more than ten thousand DNS queries per month. We then deploy a honeypot for our registered domains and collect 5,925,311 incoming queries for 6 months, from which we discover that 5,186,858 and 505,238 queries are generated from automated processes and web crawlers, respectively. Finally, we perform extensive traffic analysis on our collected data and reveal that NXDomains can be misused for various purposes, including botnet takeover, malicious file injection, and residue trust exploitation. Guannan Liu 0003, Lin Jin, Shuai Hao 0001, Yubao Zhang, Daiping Liu, Angelos Stavrou, Haining Wang 0001 |
IMC | 4 |
| 2023 | Pan-cancer analysis of SYNGR2 with a focus on clinical implications and immune landscape in liver hepatocellular carcinomaabstractBACKGROUND: Synaptogyrin-2 (SYNGR2), as a member of synaptogyrin gene family, is overexpressed in several types of cancer. However, the role of SYNGR2 in pan-cancer is largely unexplored. METHODS: From the TCGA and GEO databases, we obtained bulk transcriptomes, and clinical information. We examined the expression patterns, prognostic values, and diagnostic value of SYNGR2 in pan-cancer, and investigated the relationship of SYNGR2 expression with tumor mutation burden (TMB), microsatellite instability (MSI), immune infiltration, and immune checkpoint (ICP) genes. The gene set enrichment analysis (GSEA) software was used to perform pathway analysis. Besides, we built a nomogram of liver hepatocellular carcinoma patients (LIHC) and validated its prediction accuracy. RESULTS: SYNGR2 was highly expressed in most cancers. The high expression of SYNGR2 significantly reduced the overall survival (OS), disease-specific survival (DSS), disease-free interval (DFI), and progression-free interval (PFI) in multiple types of cancer. Also, receiver operating characteristic (ROC) curve analysis demonstrated that SYNGR2 showed high accuracy in distinguishing cancerous tissues from normal ones. Moreover, SYNGR2 expression was correlated with TMB, MSI, immune scores, and immune cell infiltrations. We also analyzed the association of SYNGR2 with immunotherapy response in LIHC. Finally, a nomogram including SYNGR2 and pathologic T, N, M stage was built and exhibited good predictive power for the OS, DSS, and PFI of LIHC patients. CONCLUSION: Overall, SYNGR2 is a critical oncogene in various tumors. SYNGR2 participates in the carcinogenic progression, and may contribute to the immune infiltration in tumor microenvironment. Our study suggests that SYNGR2 can serve as a predictor related to prognosis in pan-cancer, especially LIHC. Chunxun Liu, Zhaowei Qu, Chao Zhan, Yubao Zhang |
BMC Bioinform. | 6 |
| 2021 | Mingling of Clear and Muddy Water: Understanding and Detecting Semantic Confusion in Blackhat SEO
Kun Du, Yubao Zhang, Shuai Hao 0001, Haining Wang 0001, Jia Zhang 0004, Hai-Xin Duan |
ESORICS (1) | 3 |
| 2021 | Detecting incentivized review groups with co-review graphabstractOnline reviews play a crucial role in the ecosystem of nowadays business (especially e-commerce platforms), and have become the primary source of consumer opinions. To manipulate consumers’ opinions, some sellers of e-commerce platforms outsource opinion spamming with incentives (e.g., free products) in exchange for incentivized reviews. As incentives, by nature, are likely to drive more biased reviews or even fake reviews. Despite e-commerce platforms such as Amazon have taken initiatives to squash the incentivized review practice, sellers turn to various social networking platforms (e.g., Facebook) to outsource the incentivized reviews. The aggregation of sellers who request incentivized reviews and reviewers who seek incentives forms incentivized review groups. In this paper, we focus on the incentivized review groups in e-commerce platforms. We perform the data collections from various social networking platforms, including Facebook, WeChat, and Douban. A measurement study of incentivized review groups is conducted with regards to group members, group activities, and products. To identify the incentivized review groups, we propose a new detection approach based on co-review graphs. Specifically, we employ the community detection method to find the suspicious communities from co-review graphs. We also build a “gold standard” dataset from the data we collected, which contains the information of reviewers who belong to incentivized review groups. We utilize the “gold standard” dataset to evaluate the effectiveness of our detection approach. Yubao Zhang, Shuai Hao 0001, Haining Wang 0001 |
High Confid. Comput. | 1 |
| 2020 | Understanding Promotion-as-a-Service on GitHubabstractAs the world’s leading software development platform, GitHub has become a social networking site for programmers and recruiters who leverage its social features, such as star and fork, for career and business development. However, in this paper, we found a group of GitHub accounts that conducted promotion services in GitHub, called “promoters”, by performing paid star and fork operations on specified repositories. We also uncovered a stealthy way of tampering with historical commits, through which these promoters are able to fake commits retroactively. By exploiting such a promotion service, any GitHub user can pretend to be a skillful developer with high influence. Kun Du, Yubao Zhang, Hai-Xin Duan, Haining Wang 0001, Shuang Hao 0001, Zhou Li 0001, Min Yang 0002 |
ACSAC | 3 |
| 2020 | Review Trade: Everything Is Free in Incentivized Review Groups
Yubao Zhang, Shuai Hao 0001, Haining Wang 0001 |
SecureComm (1) | 1 |
| 2020 | Understanding the Manipulation on Recommender Systems through Web InjectionabstractRecommender systems have been increasingly used in a variety of web services, providing a list of recommended items in which a user may have an interest. While important, recommender systems are vulnerable to various malicious attacks. In this paper, we study a new security vulnerability in recommender systems caused byweb injection, through which malicious actors stealthily tamper any unprotected in-transit HTTP webpage content and force victims to visit specific items in some web services (even running HTTPS),e.g., YouTube. By doing so, malicious actors can promote their targeted items in those web services. To obtain a deeper understanding on the recommender systems of our interest (including YouTube, Yelp, Taobao, and 360 App market), we first conduct a measurement-based analysis on several real-world recommender systems by leveraging machine learning algorithms. Then, web injection is implemented in three different types of devices (i.e., computer, router, and proxy server) to investigate the scenarios where web injection could occur. Based on the implementation of web injection, we demonstrate that it is feasible and sometimes effective to manipulate the real-world recommender systems through web injection. We also present several countermeasures against such manipulations. Yubao Zhang, Jidong Xiao, Shuai Hao 0001, Haining Wang 0001, Sencun Zhu, Sushil Jajodia |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2019 | Casino royale: a deep exploration of illegal online gamblingabstractThe popularity of online gambling could bring negative social impact, and many countries ban or restrict online gambling. Taking China for example, online gambling violates Chinese laws and hence is illegal. However, illegal online gambling websites are still thriving despite strict restrictions, since they are able to make tremendous illicit profits by trapping and cheating online players. In this paper, we conduct the first deep analysis on illegal online gambling targeting Chinese to unveil its profit chain. After successfully identifying more than 967,954 suspicious illegal gambling websites, we inspect these illegal gambling websites from five aspects, including webpage structure similarity, SEO (Search Engine Optimization) methods, the abuse of Internet infrastructure, third-party online payment, and gambling group. Then we conduct a measurement study on the profit chain of illegal online gambling, investigating the upstream and downstream of these illegal gambling websites. We mainly focus on promotion strategies, third-party online payment, the abuse of third-party live chat services, and network infrastructures. Our findings shed the light on the ecosystem of online gambling and help the security community thwart illegal online gambling. Kun Du, Yubao Zhang, Shuang Hao 0001, Zhou Li 0001, Mingxuan Liu 0006, Haining Wang 0001, Hai-Xin Duan, Yazhou Shi, XiaoDong Su, Zhifeng Geng |
ACSAC | 3 |
| 2018 | End-Users Get Maneuvered: Empirical Analysis of Redirection Hijacking in Content Delivery Networks
Shuai Hao 0001, Yubao Zhang, Haining Wang 0001, Angelos Stavrou |
USENIX Security Symposium | 2 |
| 2017 | Twitter Trends Manipulation: A First Look Inside the Security of Twitter TrendingabstractTwitter trends, a timely updated set of top terms in Twitter, have the ability to affect the public agenda of the community and have attracted much attention. Unfortunately, in the wrong hands, Twitter trends can also be abused to mislead people. In this paper, we attempt to investigate whether Twitter trends are secure from the manipulation of malicious users. We collect more than 69 million tweets from 5 million accounts. Using the collected tweets, we first conduct a data analysis and discover evidence of Twitter trend manipulation. Then, we study at the topic level and infer the key factors that can determine whether a topic starts trending due to its popularity, coverage, transmission, potential coverage, or reputation. What we find is that except for transmission, all of factors above are closely related to trending. Finally, we further investigate the trending manipulation from the perspective of compromised and fake accounts and discuss countermeasures. Yubao Zhang, Xin Ruan, Haining Wang 0001, Hui Wang 0030, Su He |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2014 | What scale of audience a campaign can reach in what price on Twitter?abstractCampaigns with commercial and spam purposes have flooded the Twitter community. To understand what scale of audience a campaign could reach, we first perform a measurement study by collecting a dataset of about 10 million tweets via streaming API and one million search tweets for targeting topics, as well as 37,313 user accounts that are suspended by Twitter. From the dataset, we extract a spam campaign and a commercial promotion campaign accompanied by spamming activities. Then, we characterize the way in which a campaign can reach its audience, especially revealing the features that dominate the information diffusion. After identifying the accounts suspended by Twitter, we further inspect to what extent these features can help to weed out spam accounts. Also, the retrospective inspection is useful to uncover the tactics that malicious accounts utilize to avoid being suspended. Using the measurement results, we then develop a theoretical framework based on an epidemic model to investigate the dynamics of spammers and victims whom spammers reach in the spam campaign. With the theoretical framework, we conduct a benefit-cost analysis of the spam campaign, shedding lights on how to restrict the benefit of the spam campaign. Yubao Zhang, Xin Ruan, Haining Wang 0001, Hui Wang 0030 |
INFOCOM | 1 |
| 2011 | A geographic analysis of P2P-TV viewershipabstractA promising P2P application, P2P-TV, has attracted hundreds of thousands of Chinese viewers. These viewers who are located in different regions represent groups with distinct cultures. However, little existing research has provided sufficient insights into the societal impact of P2P-TV systems, from the viewpoint of geographic distribution of viewers. In this paper, we analyze geographic distribution of viewers of three most popular P2P-TV systems simultaneously, PPLive, PPStream and UUSee. With more than 20 GB worth of log data from three different P2P-TV systems, we have completed a thorough investigation of geographic distribution of viewers. We also seek to explore the potential correlation between viewer population density and economic development level and find that there is indeed a highly negative correlation between them. Zhihong Jiang, Hui Wang 0030, Yubao Zhang, Pei Li 0001 |
ISI | 3 |
| 2010 | The Benefits of Network Coding in Distributed Caching in Large-Scale P2P-VoD SystemsabstractDistributed caching mechanism plays an important role to improve the performance of large-scale peer-to-peer video-on-demand (P2P-VoD) systems, especially in terms of server bandwidth costs. Nevertheless, existing research and analytical studies of P2P-VoD systems have not thoroughly investigated and understood distributed caching policies and their critical properties for helping to mitigate the bandwidth costs on streaming servers. In particular, there exists no prior analytical work that focuses on a new way of designing a distributed caching strategy, with the help of network coding. In this paper, we seek to show an analytical understanding of the potential fundamental benefits of using network coding in distributed passive caching. With our problem formulation, we present probability-based expressions for computing the steady-state average server bandwidth costs, with or without the use of network coding. Our analytical results are cross-validated by our extensive simulation studies in large-scale static and dynamic scenarios. Hui Wang 0030, Yubao Zhang, Pei Li 0001, Zhihong Jiang |
GLOBECOM | 2 |