Yaowei Yan

dblp:67/8039 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
3since 2021 · last 2023
0000-0003-2359-8319ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 9 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 3 · 1 first-authorComputer networks · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2023 Random Walk on Multiple Networks
abstract
Random Walk is a basic algorithm to explore the structure of networks, which can be used in many tasks, such as local community detection and network embedding. Existing random walk methods are based on single networks that contain limited information. In contrast, real data often contain entities with different types or/and from different sources, which are comprehensive and can be better modeled by multiple networks. To take the advantage of rich information in multiple networks and make better inferences on entities, in this study, we propose random walk on multiple networks, RWM. RWM is flexible and supports both multiplex networks and general multiple networks, which may form many-to-many node mappings between networks. RWM sends a random walker on each network to obtain the local proximity (i.e., node visiting probabilities) w.r.t. the starting nodes. Walkers with similar visiting probabilities reinforce each other. We theoretically analyze the convergence properties of RWM. Two approximation methods with theoretical performance guarantees are proposed for efficient computation. We apply RWM in link prediction, network embedding, and local community detection. Comprehensive experiments conducted on both synthetic and real-world datasets demonstrate the effectiveness and efficiency of RWM.
Yuchen Bian, Yaowei Yan, Xiong Bill Yu, Jun Huan, Xiao Liu 0039, Xiang Zhang 0001
IEEE Trans. Knowl. Data Eng.3
2022 A Collective Approach to Scholar Name Disambiguation
abstract
Scholar name disambiguation remains a hard and unsolved problem, which brings various troubles for bibliography data analytics. Most existing methods handle name disambiguation separately that tackles one name at a time, and neglect the fact that disambiguation of one name affects the others. Further, it is typically common that only limited information is available for bibliography data, e.g., only basic paper and citation information is available in DBLP. In this study, we propose a collective approach to name disambiguation, which takes the connection of different ambiguous names into consideration. We reformulate bibliography data as a heterogeneous multipartite network, which initially treats each author reference as a unique author entity, and disambiguation results of one name propagate to the others of the network. To further deal with the sparsity problem caused by limited available information, we also introduce word-word and venue-venue similarities, and we finally measure author similarities by assembling similarities from four perspectives. Using real-life data, we experimentally demonstrate that our approach is both effective and efficient.
Shuai Ma 0001, Yaowei Yan, Chunming Hu, Xiang Zhang 0001, Jinpeng Huai
IEEE Trans. Knowl. Data Eng.3
2021 A Collective Approach to Scholar Name Disambiguation (Extended Abstract)
abstract
This study investigates name disambiguation for scholarly data. We propose a collective approach, which considers the connections of different ambiguous names, such that it initially treats each author reference as a unique author entity and reformulates the bibliography data as a heterogeneous multipartite network. Disambiguation results of one author name propagate to the others in the network. To further deal with the sparsity problem caused by limited available information, we also introduce word-word and venue-venue similarities and measure author similarities by assembling similarities from multiple perspectives. Using three real-life datasets, we experimentally show that our approach is both effective and efficient.
Shuai Ma 0001, Yaowei Yan, Chunming Hu, Xiang Zhang 0001, Jinpeng Huai
ICDE3
2020 Local Community Detection in Multiple Networks
abstract
Local community detection aims to find a set of densely-connected nodes containing given query nodes. Most existing local community detection methods are designed for a single network. However, a single network can be noisy and incomplete. Multiple networks are more informative in real-world applications. There are multiple types of nodes and multiple types of node proximities. Complementary information from different networks helps to improve detection accuracy. In this paper, we propose a novel RWM (Random Walk in Multiple networks) model to find relevant local communities in all networks for a given query node set from one network. RWM sends a random walker in each network to obtain the local proximity w.r.t. the query nodes (i.e., node visiting probabilities).
Yuchen Bian, Yaowei Yan, Xiao Liu 0039, Jun Huan, Xiang Zhang 0001
KDD3
2020 Memory-based random walk for multi-query local community detection
Yuchen Bian, Yaowei Yan, Wei Cheng 0002, Wei Wang 0010, Xiang Zhang 0001
Knowl. Inf. Syst.3
2020 Correction to: Memory-based random walk for multi-query local community detection
Yuchen Bian, Yaowei Yan, Wei Cheng 0002, Wei Wang 0010, Xiang Zhang 0001
Knowl. Inf. Syst.3
2019 Constrained Local Graph Clustering by Colored Random Walk
abstract
Detecting local graph clusters is an important problem in big graph analysis. Given seed nodes in a graph, local clustering aims at finding subgraphs around the seed nodes, which consist of nodes highly relevant to the seed nodes. However, existing local clustering methods either allow only a single seed node, or assume all seed nodes are from the same cluster, which is not true in many real applications. Moreover, the assumption that all seed nodes are in a single cluster fails to use the crucial information of relations between seed nodes. In this paper, we propose a method to take advantage of such relationship. With prior knowledge of the community membership of the seed nodes, the method labels seed nodes in the same (different) community by the same (different) color. To further use this information, we introduce a color-based random walk mechanism, where colors are propagated from the seed nodes to every node in the graph. By the interaction of identical and distinct colors, we can enclose the supervision of seed nodes into the random walk process. We also propose a heuristic strategy to speed up the algorithm by more than 2 orders of magnitude. Experimental evaluations reveal that our clustering method outperforms state-of-the-art approaches by a large margin.
Yaowei Yan, Yuchen Bian, Dongwon Lee 0001, Xiang Zhang 0001
WWW1
2018 On Multi-query Local Community Detection
abstract
Local community detection, which aims to find a target community containing a set of query nodes, has recently drawn intense research interest. The existing local community detection methods usually assume all query nodes are from the same community and only find a single target community. This is a strict requirement and does not allow much flexibility. In many real-world applications, however, we may not have any prior knowledge about the community memberships of the query nodes, and different query nodes may be from different communities. To address this limitation of the existing methods, we propose a novel memory-based random walk method, MRW, that can simultaneously identify multiple target local communities to which the query nodes belong. In MRW, each query node is associated with a random walker. Different from commonly used memoryless random walk models, MRW records the entire visiting history of each walker. The visiting histories of walkers can help unravel whether they are from the same community or not. Intuitively, walkers with similar visiting histories are more likely to be in the same community. Moreover, MRW allows walkers with similar visiting histories to reinforce each other so that they can better capture the community structure instead of being biased to the query nodes. We provide rigorous theoretical foundation for the proposed method and develop efficient algorithms to identify multiple target local communities simultaneously. Comprehensive experimental evaluations on a variety of real-world datasets demonstrate the effectiveness and efficiency of the proposed method.
Yuchen Bian, Yaowei Yan, Wei Cheng 0002, Wei Wang 0010, Xiang Zhang 0001
ICDM2
2018 Local Graph Clustering by Multi-network Random Walk with Restart
Yaowei Yan, Jingchao Ni, Hongliang Fei, Wei Fan 0001, Xiong Bill Yu, John Yen, Xiang Zhang 0001
PAKDD (3)1
2015 Accelerating SAT Solving by Common Subclause Elimination
abstract
Boolean SATisfiability (SAT) is an important problem in AI. SAT solvers have been effectively used in important industrial applications including automated planning and verification. In this paper, we present novel algorithms for fast SAT solving by employing two common subclause elimination (CSE) approaches. Our motivation is that modern SAT solving techniques can be more efficient on CSE-processed instances. Empirical study shows that CSE can significantly speed up SAT solving.
Yaowei Yan, Chris E. Gutierrez, Jeriah Jn-Charles, Forrest Sheng Bao, Yuanlin Zhang 0002
AAAI1
2015 Gossiping along the Path: A Direction-Biased Routing Scheme for Wireless Ad Hoc Networks
abstract
Like any communication networks, Wireless ad hoc networks (WANETs) require routing to support many applications such as data aggregation and over-the- air firmware update. Traditional route discovery in WANETs floods request all over the network causing broadcast storm that will heavily consume the precious power resources on nodes. To respond, gossip routing is proposed to reduce the number of messages. In recent years, several algorithms have been developed to improve its efficiency using location information. However, they only make use of distance but not directional information which is much cheaper to acquire in WANETs. In this paper, we develop an approach to improve location- aided routing by making use of directional information, with and without distance information. Empirical results show that our approach outperforms existing gossip routing algorithms, including location-aided ones, on random and deployed WSNs. When only general direction information of nodes is given, performances of our algorithms only drop slightly. Hoping this approach can also benefit other types of routing tasks, we also migrate the spirit of this algorithm onto opportunistic routing and show improved performance than existing location-aided opportunistic routing.
Yaowei Yan, Nghi Huu Tran, Forrest Sheng Bao
GLOBECOM1
2009 vBench: A micro-benchmark for File - I/O performance of virtual machines
abstract
Resurgence of virtualization calls for suitable benchmark to validate the design of virtualization system. Among the virtual machines' performance, I/O especially disk I/O is the most critical one. Since I/O is a key factor that impacts virtual machine performance, this paper introduces a benchmark- vBench for disk I/O micro-performance evaluation of virtual machine. vBench is mainly consisted of serial disk I/O evaluation and parallel disk I/O evaluation. Since error during test is inevitable, for the purpose of error correction, we design an efficient approach which we runs the same workloads in two loops, then the execution times of two loops minus each other. The approach can correctly deal with loop overhead. To prevent compiler from optimizing the code and lead to the wrong execution time of workloads, vBench is embedded inline assembly code. vBench provides the performance evaluation for system call of disk I/O, application level disk I/O and parallel disk I/O. Disk I/O performance evaluation of vBench can measure disk I/O micro-performance of virtual machine in certain extreme conditions. Experiments show that the results of vBench are accurate and vBench can reflect the disk I/O performance of the virtual machines correctly. Therefore, vBench is more accurate, more reliable than the traditional benchmarks and more suitable for evaluating virtual machine disk I/O performance.
Pingpeng Yuan, Hai Jin 0001, Ding Ye, Wenzhi Cao, Yaowei Yan
APSCC5