Yubao Wu

dblp:18/8107 · DBLP profile ↗
← Back
20ranked-venue papers
9as first author
4since 2021 · last 2025
0000-0001-9356-8508ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 12 · 7 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 4 · 2 since 2021Computer networks · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
6 papers
Graph data management · 47% Data mining · 45% Spatial and temporal data management · 8%
Theoretical computer science
5 papers
Graph algorithms and graph theory · 54% Mathematical optimization · 28% Algorithms and data structures · 14%
Artificial intelligence
1 paper
Graph learning · 100%

Topics — the 23 heaviest of 24, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Graph data management › graph algorithms
second-order random walk
0.622018
Second-order random walk-based proximity measures in graph analysis: formulations and algorithms · VLDB J. 2018
Remember Where You Came From: On The Second-Order Random Walk Based Proximity Measures · Proc. VLDB Endow. 2016
Mathematical optimization › combinatorial optimization
local search
0.422016
Efficient and Exact Local Search for Random Walk Based Top-K Proximity Query in Large Graphs · IEEE Trans. Knowl. Data Eng. 2016
Fast and unified local search for random walk based k-nearest-neighbor query in large graphs · SIGMOD Conference 2014
Graph data management
graph analytics
0.312018
Second-order random walk-based proximity measures in graph analysis: formulations and algorithms · VLDB J. 2018
Graph data management › graph representation
proximity graph
0.212016
Remember Where You Came From: On The Second-Order Random Walk Based Proximity Measures · Proc. VLDB Endow. 2016
Spatial and temporal data management › spatial query processing
proximity search
0.212016
Efficient and Exact Local Search for Random Walk Based Top-K Proximity Query in Large Graphs · IEEE Trans. Knowl. Data Eng. 2016
Graph algorithms and graph theory
random walk
0.212016
Efficient and Exact Local Search for Random Walk Based Top-K Proximity Query in Large Graphs · IEEE Trans. Knowl. Data Eng. 2016
Data mining › structured data mining › graph mining
community detection
0.212015
Robust Local Community Detection: On Free Rider Effect and Its Elimination · Proc. VLDB Endow. 2015
Graph data management
dense subgraph
0.212015
Finding dense and connected subgraphs in dual networks · ICDE 2015
Data mining › structured data mining
graph mining
0.212015
Robust Local Community Detection: On Free Rider Effect and Its Elimination · Proc. VLDB Endow. 2015
Data mining › pattern mining
graph pattern mining
0.212015
Finding dense and connected subgraphs in dual networks · ICDE 2015
Data mining › structured data mining › graph mining › community detection
local community detection
0.212015
Robust Local Community Detection: On Free Rider Effect and Its Elimination · Proc. VLDB Endow. 2015
Graph algorithms and graph theory › dense subgraph discovery
densest subgraph
0.212015
Robust Local Community Detection: On Free Rider Effect and Its Elimination · Proc. VLDB Endow. 2015
Graph algorithms and graph theory
dense subgraph discovery
0.212015
Finding dense and connected subgraphs in dual networks · ICDE 2015
Algorithms and data structures › similarity search › nearest neighbor search
k-nearest neighbors
0.212014
Fast and unified local search for random walk based k-nearest-neighbor query in large graphs · SIGMOD Conference 2014
Graph algorithms and graph theory › random walk
random-walk similarity
0.212014
Fast and unified local search for random walk based k-nearest-neighbor query in large graphs · SIGMOD Conference 2014
Machine learning › Graph learning
graph clustering
0.212013
Flexible and robust co-regularized multi-domain graph clustering · KDD 2013
Data mining
clustering
0.212013
Flexible and robust co-regularized multi-domain graph clustering · KDD 2013
Data mining › clustering › graph clustering
multi-view graph clustering
0.212013
Flexible and robust co-regularized multi-domain graph clustering · KDD 2013
Graph algorithms and graph theory
network analysis
0.112016
Remember Where You Came From: On The Second-Order Random Walk Based Proximity Measures · Proc. VLDB Endow. 2016
Computational geometry
proximity measures
0.112016
Efficient and Exact Local Search for Random Walk Based Top-K Proximity Query in Large Graphs · IEEE Trans. Knowl. Data Eng. 2016
Data mining
pattern mining
0.112015
Robust Local Community Detection: On Free Rider Effect and Its Elimination · Proc. VLDB Endow. 2015
Data mining › structured data mining › graph mining
subgraph mining
0.112015
Robust Local Community Detection: On Free Rider Effect and Its Elimination · Proc. VLDB Endow. 2015
Mathematical optimization › integer programming
proximity bounds
0.112014
Fast and unified local search for random walk based k-nearest-neighbor query in large graphs · SIGMOD Conference 2014

Methods — techniques the papers use, named apart from their topics

lower and upper bounds · 0.5local search · 0.5incidence matrix representation · 0.5pruning · 0.4node weighting · 0.4greedy method · 0.4density metric · 0.4random walk · 0.3co-regularization · 0.3monte carlo methods · 0.2monte carlo method · 0.2lower and upper bound computation · 0.2non-negative matrix factorization · 0.2
YearPublicationVenuePosition
2025 Towards Explainable and Educational Phishing Detection: A Zero-Shot LLM Approach
Yubao Wu
IEEE Big Data2
2022 Computational Approaches to Detect Illicit Drug Ads and Find Vendor Communities Within Social Media Platforms
abstract
The opioid abuse epidemic represents a major public health threat to global populations. The role social media may play in facilitating illicit drug trade is largely unknown due to limited research. However, it is known that social media use among adults in the US is widespread, there is vast capability for online promotion of illegal drugs with delayed or limited deterrence of such messaging, and further, general commercial sale applications provide safeguards for transactions; however, they do not discriminate between legal and illegal sale transactions. These characteristics of the social media environment present challenges to surveillance which is needed for advancing knowledge of online drug markets and the role they play in the drug abuse and overdose deaths. In this paper, we present a computational framework developed to automatically detect illicit drug ads and communities of vendors. The SVM- and CNN- based methods for detecting illicit drug ads, and a matrix factorization based method for discovering overlapping communities have been extensively validated on the large dataset collected from Google+, Flickr and Tumblr. Pilot test results demonstrate that our computational methods can effectively identify illicit drug ads and detect vendor-community with accuracy. These methods hold promise to advance scientific knowledge surrounding the role social media may play in perpetuating the drug abuse epidemic.
Fengpan Zhao, Pavel Skums, Alex Zelikovsky, Eric L. Sevigny, Monica Haavisto Swahn, Sheryl M. Strasser, Yan Huang 0032, Yubao Wu
IEEE ACM Trans. Comput. Biol. Bioinform.8
2021 Hidden Buyer Identification in Darknet Markets via Dirichlet Hawkes Process
abstract
Darknet markets are underground markets for various illicit transactions, including selling or brokering drugs, weapons, and stolen credit cards. To combat these illicit activities in cyberspace, it is critical to understand the activity behaviors of participants in the darknet markets. Currently, many studies focus on studying the activities of vendors. However, there is no much work on analyzing buyers. The key challenge is that the buyers are anonymized in darknet markets. To ensure the anonymity of transactions, we only observe the first a nd last digits of a buyer’s ID, such as "a**b", on most of the darknet markets. To tackle this challenge, we propose a hidden buyer identification model, called UNMIX, which can group transactions from one hidden buyer into one cluster given a transaction sequence from an anonymized ID. UNMIX is able to model the temporal dynamics information as well as the product, comment, and vendor information associated with each transaction. Then, the transactions with similar patterns in terms of time and content are grouped as a subsequence from one hidden buyer. Experiments on the data collected from three real-world darknet markets and one DBLP publication dataset demonstrate the effectiveness of our approach measured by various clustering metrics. Case studies on real transaction sequences explicitly show that our approach can group transactions with similar patterns into the same clusters.
Panpan Zheng, Shuhan Yuan, Xintao Wu, Yubao Wu
IEEE BigData4
2021 Should We Trust Influencers on Social Networks? On Instagram Sponsored Post Analysis
abstract
With online social networks (OSNs), people are exposed to tons of fake information or misleading posts. Celebrities sometimes intentionally create misleading posts in OSNs to guide people for commercial or marketing purposes. The intentional phrases from such posts can affect the online rating and even lead to a frenzy shopping. That’s part of the reasons that the top social media influencers are targeted by merchants to help promote products. The Federal Trade Commission (FTC) requires that all sponsored posts must be clearly disclosed. However, many influencers do not follow the FTC rules. As a result, people may be misled by the undisclosed sponsorship. In this study, for the first time, we explore the credibility of posts on Instagram and analyze if an influencer complies with the FTC requirements. We build an effective Undisclosed Sponsored Post Detection (USPD) framework based on an ensemble of machine learning classifiers. The USPD framework consists of three main processes: (i) feature extraction, (ii) model construction and (iii) credibility and integrity analysis. Our analysis and experiments demonstrate that the proposed framework can achieve a high accuracy of 83% for undisclosed sponsored post detection. The proposed framework also takes advantages of the text, user and image features in OSN posts to effectively analyze how much an influencer can be trusted.
Xueting Liao, Danyang Zheng 0001, Yubao Wu, Xiaojun Cao
ICCCN3
2020 Towards k-vertex connected component discovery from large networks
Yuan Li 0008, Guoren Wang, Yuhai Zhao, Feida Zhu 0001, Yubao Wu
World Wide Web5
2019 Into the Reverie: Exploration of the Dream Market
abstract
Since the emergence of the Silk Road market in the early 2010s, dark web `cryptomarkets' have proliferated and offered people an online platform to buy and sell illicit drugs, relying on cryptocurrencies such as Bitcoin for anonymous transactions. However, recent studies have highlighted the potential for de-anonymization of bitcoin transactions, bringing into question the level of anonymity afforded by cryptomarkets. We examine a set of over 100,000 product reviews from several cryptomarkets collected in 2018 and 2019 and conduct a comprehensive analysis of the markets, including an examination of the distribution of drug sales and revenue among vendors, and a comparison of incidences of opioid sales to overdose deaths in a US city. We explore the potential for de-anonymization of vendors by implementing a Naïve-Bayes classifier to predict the vendor from a given product review, and attempt to link vendors' sales to specific Bitcoin transactions. On the buyer side, we evaluate the efficacy of hierarchical agglomerative clustering for grouping together transactions corresponding to the same buyer. We find that the high degree of specialization among the small subset of high-revenue vendors may render these vendors susceptible to de-anonymization. Further research is necessary to confirm these findings, which are restricted by the scarcity of ground-truth data for validation.
Theo Carr, Jun Zhuang 0004, Dwight Sablan, Emma LaRue, Yubao Wu, Mohammad Al Hasan, George O. Mohler
IEEE BigData5
2019 Second-Order CoSimRank for Similarity Measures in Social Networks
abstract
Measuring the similarity between nodes is challenging in social networks. The SimRank and CoSimRank are techniques widely used to calculate the similarity of two nodes in a social graph. They can be applied to many applications such as recommending friends and detecting communities in social networks. Both SimRank and CoSimRank are based on random walk and only consider first-order transition probabilities, in which the next node to visit in random walk solely depends on the current node, like a Markov chain. However, in many real-world situations, simply considering the current node may not be enough. Previously visited node may provide extra information for measuring similarities. In this paper, we propose a novel similarity measure technique by investigating CoSimRank to take advantage of the second-order information in a random walk process. Our extensive analysis and experiments show that the proposed second-order CoSimRank significantly outperforms the existing techniques.
Xueting Liao, Yubao Wu, Xiaojun Cao
ICC2
2019 Detecting Illicit Drug Ads in Google+ Using Machine Learning
Fengpan Zhao, Pavel Skums, Alex Zelikovsky, Eric L. Sevigny, Monica Haavisto Swahn, Sheryl M. Strasser, Yubao Wu
ISBRA7
2019 A Second-Order Diffusion Model for Influence Maximization in Social Networks
abstract
In social networks, several influential individuals can promote an idea or a product to numerous individuals. Thus, it is valuable to solve the influence maximization (IM) problem, which asks for finding the most influential set of individuals in a social network. To estimate the influence of individuals, the existing independent cascade (IC) model simulates the influence diffusion only considering the influences from direct in-neighbors to nodes. This consideration does not hold in real life. In many cases, people are likely influenced by information depending on where it comes from, instead of who gives it. To simulate the influence diffusion more accurate, this paper proposes the second-order IC model, which takes the previous influence into consideration. In addition, we design an approximate algorithm and its distributed extension for IM under the second-order IC model. Experimental results show that our second-order IC model outperforms the IC model in terms of simulating influence diffusions. The proposed algorithms are efficient, and the obtained node sets are influential.
Wenyi Tang, Guangchun Luo, Yubao Wu, Ling Tian, Xu Zheng 0001, Zhipeng Cai 0001
IEEE Trans. Comput. Soc. Syst.3
2018 Predicting Opioid Epidemic by Using Twitter Data
Yubao Wu, Pavel Skums, Alex Zelikovsky, David S. Campo, Xueting Liao
ISBRA1
2018 Second-order random walk-based proximity measures in graph analysis: formulations and algorithms
Yubao Wu, Xiang Zhang 0001, Yuchen Bian, Zhipeng Cai 0001, Xiang Lian 0001, Xueting Liao, Fengpan Zhao
VLDB J.1
2017 Effective k-Vertex Connected Component Detection in Large-Scale Networks
Yuan Li 0008, Yuhai Zhao, Guoren Wang, Feida Zhu 0001, Yubao Wu, Shengle Shi
DASFAA (2)5
2016 Remember Where You Came From: On The Second-Order Random Walk Based Proximity Measures
abstract
Measuring the proximity between different nodes is a fundamental problem in graph analysis. Random walk based proximity measures have been shown to be effective and widely used. Most existing random walk measures are based on the first-order Markov model, i.e., they assume that the next step of the random surfer only depends on the current node. However, this assumption neither holds in many real-life applications nor captures the clustering structure in the graph. To address the limitation of the existing first-order measures, in this paper, we study the second-order random walk measures, which take the previously visited node into consideration. While the existing first-order measures are built on node-to-node transition probabilities, in the second-order random walk, we need to consider the edge-to-edge transition probabilities. Using incidence matrices, we develop simple and elegant matrix representations for the second-order proximity measures. A desirable property of the developed measures is that they degenerate to their original first-order forms when the effect of the previous step is zero. We further develop Monte Carlo methods to efficiently compute the second-order measures and provide theoretical performance guarantees. Experimental results show that in a variety of applications, the second-order measures can dramatically improve the performance compared to their first-order counterparts.
Yubao Wu, Yuchen Bian, Xiang Zhang 0001
Proc. VLDB Endow.1
2016 Mining Dual Networks: Models, Algorithms, and Applications
abstract
Finding the densest subgraph in a single graph is a fundamental problem that has been extensively studied. In many emerging applications, there exist dual networks. For example, in genetics, it is important to use protein interactions to interpret genetic interactions. In this application, one network represents physical interactions among nodes, for example, protein--protein interactions, and another network represents conceptual interactions, for example, genetic interactions. Edges in the conceptual network are usually derived based on certain correlation measure or statistical test measuring the strength of the interaction. Two nodes with strong conceptual interaction may not have direct physical interaction. In this article, we propose the novel dual-network model and investigate the problem of finding the densest connected subgraph (DCS), which has the largest density in the conceptual network and is also connected in the physical network. Density in the conceptual network represents the average strength of the measured interacting signals among the set of nodes. Connectivity in the physical network shows how they interact physically. Such pattern cannot be identified using the existing algorithms for a single network. We show that even though finding the densest subgraph in a single network is polynomial time solvable, the DCS problem is NP-hard. We develop a two-step approach to solve the DCS problem. In the first step, we effectively prune the dual networks, while guarantee that the optimal solution is contained in the remaining networks. For the second step, we develop two efficient greedy methods based on different search strategies to find the DCS. Different variations of the DCS problem are also studied. We perform extensive experiments on a variety of real and synthetic dual networks to evaluate the effectiveness and efficiency of the developed methods.
Yubao Wu, Xiaofeng Zhu 0003, Wei Fan 0001, Ruoming Jin, Xiang Zhang 0001
ACM Trans. Knowl. Discov. Data1
2016 Efficient and Exact Local Search for Random Walk Based Top-K Proximity Query in Large Graphs
abstract
Top-$k$proximity query in large graphs is a fundamental problem with a wide range of applications. Various random walk based measures have been proposed to measure the proximity between different nodes. Although these measures are effective, efficiently computing them on large graphs is a challenging task. In this paper, we develop an efficient and exact local search method, FLoS (Fast Local Search), for top-$k$proximity query in large graphs. FLoS guarantees the exactness of the solution. Moreover, it can be applied to a variety of commonly used proximity measures. FLoS is based on theno local optimumproperty of proximity measures. We show that many measures have no local optimum. Utilizing this property, we introduce several operations to manipulate transition probabilities and develop tight lower and upper bounds on the proximity values. The lower and upper bounds monotonically converge to the exact proximity value when more nodes are visited. We further extend FLoS to measures having local optimum by utilizing relationship among different measures. We perform comprehensive experiments on real and synthetic large graphs to evaluate the efficiency and effectiveness of the proposed method.
Yubao Wu, Ruoming Jin, Xiang Zhang 0001
IEEE Trans. Knowl. Data Eng.1
2015 Finding dense and connected subgraphs in dual networks
abstract
Finding dense subgraphs is an important problem that has recently attracted a lot of interests. Most of the existing work focuses on a single graph (or network1). In many real-life applications, however, there exist dual networks, in which one network represents the physical world and another network represents the conceptual world. In this paper, we investigate the problem of finding the densest connected subgraph (DCS) which has the largest density in the conceptual network and is also connected in the physical network. Such pattern cannot be identified using the existing algorithms for a single network. We show that even though finding the densest subgraph in a single network is polynomial time solvable, the DCS problem is NP-hard. We develop a two-step approach to solve the DCS problem. In the first step, we effectively prune the dual networks while guarantee that the optimal solution is contained in the remaining networks. For the second step, we develop two efficient greedy methods based on different search strategies to find the DCS. Different variations of the DCS problem are also studied. We perform extensive experiments on a variety of real and synthetic dual networks to evaluate the effectiveness and efficiency of the developed methods.
Yubao Wu, Ruoming Jin, Xiaofeng Zhu 0003, Xiang Zhang 0001
ICDE1
2015 Robust Local Community Detection: On Free Rider Effect and Its Elimination
abstract
Given a large network, local community detection aims at finding the community that contains a set of query nodes and also maximizes (minimizes) a goodness metric. This problem has recently drawn intense research interest. Various goodness metrics have been proposed. However, most existing metrics tend to include irrelevant subgraphs in the detected local community. We refer to such irrelevant subgraphs as free riders. We systematically study the existing goodness metrics and provide theoretical explanations on why they may cause the free rider effect. We further develop a query biased node weighting scheme to reduce the free rider effect. In particular, each node is weighted by its proximity to the query node. We define a query biased density metric to integrate the edge and node weights. The query biased densest subgraph, which has the largest query biased density, will shift to the neighborhood of the query nodes after node weighting. We then formulate the query biased densest connected subgraph (QDC) problem, study its complexity, and provide efficient algorithms to solve it. We perform extensive experiments on a variety of real and synthetic networks to evaluate the effectiveness and efficiency of the proposed methods.
Yubao Wu, Ruoming Jin, Jing Li 0002, Xiang Zhang 0001
Proc. VLDB Endow.1
2014 Fast and unified local search for random walk based k-nearest-neighbor query in large graphs
abstract
Given a large graph and a query node, finding its k-nearest-neighbor (kNN) is a fundamental problem. Various random walk based measures have been developed to measure the proximity (similarity) between nodes. Existing algorithms for the random walk based top-k proximity search can be categorized as global and local methods based on their search strategies. Global methods usually require an expensive precomputing step. By only searching the nodes near the query node, local methods have the potential to support more efficient query. However, most existing local search methods cannot guarantee the exactness of the solution. Moreover, they are usually designed for specific proximity measures. Can we devise an efficient local search method that applies to different measures and also guarantees result exactness? In this paper, we present FLoS (Fast Local Search), a unified local search method for efficient and exact top-k proximity query in large graphs. FLoS is based on the no local optimum property of proximity measures. We show that many measures have no local optimum. Utilizing this property, we introduce several simple operations on transition probabilities, which allow developing lower and upper bounds on the proximity. The bounds monotonically converge to the exact proximity when more nodes are visited. We further show that FLoS can also be applied to measures having local optimum by utilizing relationship among different measures. We perform comprehensive experiments to evaluate the efficiency and applicability of the proposed method.
Yubao Wu, Ruoming Jin, Xiang Zhang 0001
SIGMOD Conference1
2013 Flexible and robust co-regularized multi-domain graph clustering
abstract
Multi-view graph clustering aims to enhance clustering performance by integrating heterogeneous information collected in different domains. Each domain provides a different view of the data instances. Leveraging cross-domain information has been demonstrated an effective way to achieve better clustering results. Despite the previous success, existing multi-view graph clustering methods usually assume that different views are available for the same set of instances. Thus instances in different domains can be treated as having strict one-to-one relationship. In many real-life applications, however, data instances in one domain may correspond to multiple instances in another domain. Moreover, relationships between instances in different domains may be associated with weights based on prior (partial) knowledge. In this paper, we propose a flexible and robust framework, CGC (Co-regularized Graph Clustering), based on non-negative matrix factorization (NMF), to tackle these challenges. CGC has several advantages over the existing methods. First, it supports many-to-many cross-domain instance relationship. Second, it incorporates weight on cross-domain relationship. Third, it allows partial cross-domain mapping so that graphs in different domains may have different sizes. Finally, it provides users with the extent to which the cross-domain instance relationship violates the in-domain clustering structure, and thus enables users to re-evaluate the consistency of the relationship. Extensive experimental results on UCI benchmark data sets, newsgroup data sets and biological interaction networks demonstrate the effectiveness of our approach.
Wei Cheng 0002, Xiang Zhang 0001, Zhishan Guo, Yubao Wu, Patrick F. Sullivan, Wei Wang 0010
KDD4
2009 Printer forensics based on page document's geometric distortion
abstract
A printed document can provide intrinsic features of the printer so as to distinguish which printer it comes from. But how to extract the intrinsic features is critical in printer forensics. In this paper, the page document's geometric distortion is extracted as the intrinsic features, and a printer forensics method based on the distortion is proposed. Firstly projective transformation model is used to model the geometric distortion. After the feature point set of the model is extracted, the model's parameters considered as the geometric distortion features can be estimated, and then the model's error pattern can be obtained. During the process, the least squares method is used to estimate the model's parameters, and SVM technique is used for classification. The effectiveness of the model's parameters in the printer forensics is demonstrated by experimental results.
Yubao Wu, Xiangwei Kong 0001, Xingang You, Yiping Guo
ICIP1