Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Yuchen Bian

dblp:187/4068 · DBLP profile ↗
← Back
21ranked-venue papers
7as first author
9since 2021 · last 2024
0000-0002-0685-3771ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 4 first-author · 7 since 2021Databases, data management, data science and information retrieval · 12 · 6 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
8 papers
Data mining · 63% Knowledge graphs · 18% Graph data management · 16%
Artificial intelligence
5 papers
Graph learning · 34% Trustworthy machine learning · 28% Efficient and distributed learning · 18%
Theoretical computer science
7 papers
Graph algorithms and graph theory · 100%

Topics — the 27 heaviest of 28, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data mining › structured data mining
graph mining
1.742023
Random Walk on Multiple Networks · IEEE Trans. Knowl. Data Eng. 2023
Rethinking Local Community Detection: Query Nodes Replacement · ICDM 2020
On Multi-query Local Community Detection · ICDM 2018
Data mining › structured data mining › graph mining › community detection
local community detection
1.742023
Random Walk on Multiple Networks · IEEE Trans. Knowl. Data Eng. 2023
Rethinking Local Community Detection: Query Nodes Replacement · ICDM 2020
On Multi-query Local Community Detection · ICDM 2018
Knowledge graphs
link prediction
1.222023
Random Walk on Multiple Networks · IEEE Trans. Knowl. Data Eng. 2023
Data Collection vs. Knowledge Graph Completion: What is Needed to Improve Coverage? · EMNLP (1) 2021
Graph algorithms and graph theory
random walk
1.132023
Random Walk on Multiple Networks · IEEE Trans. Knowl. Data Eng. 2023
Constrained Local Graph Clustering by Colored Random Walk · WWW 2019
Many Heads are Better than One: Local Community Detection by the Multi-walker Chain · ICDM 2017
Machine learning › Graph learning
graph neural network
0.812024
Generating In-Distribution Proxy Graphs for Explaining Graph Neural Networks · ICML 2024
Machine learning › Trustworthy machine learning › interpretability
graph neural network explanation
0.812024
Generating In-Distribution Proxy Graphs for Explaining Graph Neural Networks · ICML 2024
Machine learning › Trustworthy machine learning
interpretability
0.812024
Generating In-Distribution Proxy Graphs for Explaining Graph Neural Networks · ICML 2024
Data mining › structured data mining › graph mining
network embedding
0.712023
Random Walk on Multiple Networks · IEEE Trans. Knowl. Data Eng. 2023
Graph data management › graph algorithms
second-order random walk
0.622018
Second-order random walk-based proximity measures in graph analysis: formulations and algorithms · VLDB J. 2018
Remember Where You Came From: On The Second-Order Random Walk Based Proximity Measures · Proc. VLDB Endow. 2016
Natural language and speech › Speech recognition and synthesis › automatic speech recognition › end-to-end speech recognition
connectionist temporal classification
0.612022
W-CTC: a Connectionist Temporal Classification Loss with Wild Cards · ICLR 2022
Machine learning › Efficient and distributed learning › model compression › pruning › structured pruning
channel pruning
0.512021
Automatic Channel Pruning with Hyper-parameter Search and Dynamic Masking · ACM Multimedia 2021
Natural language and speech › Language models and text generation › text representation
contextualized word embeddings
0.512021
Isotropy in the Contextual Embedding Space: Clusters and Manifolds · ICLR 2021
Machine learning › Efficient and distributed learning
model compression
0.512021
Automatic Channel Pruning with Hyper-parameter Search and Dynamic Masking · ACM Multimedia 2021
Machine learning › Graph learning
graph representation learning
0.412020
Deep Multi-Graph Clustering via Attentive Cross-Graph Association · WSDM 2020
Machine learning › Graph learning
network embedding
0.412020
Deep Multi-Graph Clustering via Attentive Cross-Graph Association · WSDM 2020
Data mining
clustering
0.412020
Deep Multi-Graph Clustering via Attentive Cross-Graph Association · WSDM 2020
Graph algorithms and graph theory › graph clustering
community detection
0.412020
Local Community Detection in Multiple Networks · KDD 2020
Graph algorithms and graph theory › graph clustering › community detection
local community detection
0.412020
Local Community Detection in Multiple Networks · KDD 2020
Graph algorithms and graph theory
graph clustering
0.412019
Constrained Local Graph Clustering by Colored Random Walk · WWW 2019
Graph algorithms and graph theory › graph clustering
local graph clustering
0.412019
Constrained Local Graph Clustering by Colored Random Walk · WWW 2019
Graph data management
graph analytics
0.312018
Second-order random walk-based proximity measures in graph analysis: formulations and algorithms · VLDB J. 2018
Graph data management › graph representation
proximity graph
0.212016
Remember Where You Came From: On The Second-Order Random Walk Based Proximity Measures · Proc. VLDB Endow. 2016
Machine learning › Graph learning
graph generation
0.212024
Generating In-Distribution Proxy Graphs for Explaining Graph Neural Networks · ICML 2024
Machine learning and data management
data collection
0.112021
Data Collection vs. Knowledge Graph Completion: What is Needed to Improve Coverage? · EMNLP (1) 2021
Knowledge graphs
knowledge graph construction
0.112021
Data Collection vs. Knowledge Graph Completion: What is Needed to Improve Coverage? · EMNLP (1) 2021
Graph algorithms and graph theory › centrality › pagerank
personalized pagerank
0.112020
Rethinking Local Community Detection: Query Nodes Replacement · ICDM 2020
Graph algorithms and graph theory
network analysis
0.112016
Remember Where You Came From: On The Second-Order Random Walk Based Proximity Measures · Proc. VLDB Endow. 2016

Methods — techniques the papers use, named apart from their topics

convergence analysis · 1.3random walk · 1.1topology potential · 0.9random walk with restart · 0.9autoencoder · 0.9attention mechanism · 0.9information theory · 0.8graph generator · 0.8visiting history · 0.7memory-based random walk · 0.7approximation methods · 0.7approximation method · 0.7hyperparameter search · 0.5dynamic masking · 0.5automated machine learning · 0.5minimum-entropy clustering · 0.4minimum entropy clustering · 0.4heuristic speedup · 0.4
YearPublicationVenuePosition
2024 Generating In-Distribution Proxy Graphs for Explaining Graph Neural Networks
abstract
Graph Neural Networks (GNNs) have become a building block in graph data processing, with wide applications in critical domains. The growing needs to deploy GNNs in high-stakes applications necessitate explainability for users in the decision-making processes. A popular paradigm for the explainability of GNNs is to identify explainable subgraphs by comparing their labels with the ones of original graphs. This task is challenging due to the substantial distributional shift from the original graphs in the training set to the set of explainable subgraphs, which prevents accurate prediction of labels with the subgraphs. To address it, in this paper, we propose a novel method that generates proxy graphs for explainable subgraphs that are in the distribution of training data. We introduce a parametric method that employs graph generators to produce proxy graphs. A new training objective based on information theory is designed to ensure that proxy graphs not only adhere to the distribution of training data but also preserve explanatory factors. Such generated proxy graphs can be reliably used to approximate the predictions of the labels of explainable subgraphs. Empirical evaluations across various datasets demonstrate our method achieves more accurate explanations for GNNs.
Zhuomin Chen, Jiaxing Zhang 0002, Jingchao Ni, Yuchen Bian, Md Mezbahul Islam, M. Mondal Ananda, Hua Wei 0001
ICML5
2023 Random Walk on Multiple Networks
abstract
Random Walk is a basic algorithm to explore the structure of networks, which can be used in many tasks, such as local community detection and network embedding. Existing random walk methods are based on single networks that contain limited information. In contrast, real data often contain entities with different types or/and from different sources, which are comprehensive and can be better modeled by multiple networks. To take the advantage of rich information in multiple networks and make better inferences on entities, in this study, we propose random walk on multiple networks, RWM. RWM is flexible and supports both multiplex networks and general multiple networks, which may form many-to-many node mappings between networks. RWM sends a random walker on each network to obtain the local proximity (i.e., node visiting probabilities) w.r.t. the starting nodes. Walkers with similar visiting probabilities reinforce each other. We theoretically analyze the convergence properties of RWM. Two approximation methods with theoretical performance guarantees are proposed for efficient computation. We apply RWM in link prediction, network embedding, and local community detection. Comprehensive experiments conducted on both synthetic and real-world datasets demonstrate the effectiveness and efficiency of RWM.
Yuchen Bian, Yaowei Yan, Xiong Bill Yu, Jun Huan, Xiao Liu 0039, Xiang Zhang 0001
IEEE Trans. Knowl. Data Eng.2
2022 W-CTC: a Connectionist Temporal Classification Loss with Wild Cards
Xingyu Cai, Jiahong Yuan, Yuchen Bian, Guangxu Xun, Jiaji Huang, Kenneth Church 0001
ICLR3
2022 Training on Lexical Resources
abstract
We propose using lexical resources (thesaurus, VAD) to fine-tune pretrained deep nets such as BERT and ERNIE. Then at inference time, these nets can be used to distinguish synonyms from antonyms, as well as VAD distances. The inference method can be applied to words as well as texts such as multiword expressions (MWEs), out of vocabulary words (OOVs), morphological variants and more. Code and data are posted on https://github.com/kwchurch/syn_ant.
Kenneth Church 0001, Xingyu Cai, Yuchen Bian
LREC3
2022 Emerging trends: General fine-tuning (gft)
abstract
Abstract This paper describes gft (general fine-tuning), a little language for deep nets, introduced at an ACL-2022 tutorial. gft makes deep nets accessible to a broad audience including non-programmers. It is standard practice in many fields to use statistics packages such as R. One should not need to know how to program in order to fit a regression or classification model and to use the model to make predictions for novel inputs. With gft, fine-tuning and inference are similar to fit and predict in regression and classification. gft demystifies deep nets; no one would suggest that regression-like methods are “intelligent.”
Kenneth Church 0001, Xingyu Cai, Yibiao Ying, Guangxu Xun, Yuchen Bian
Nat. Lang. Eng.6
2021 Data Collection vs. Knowledge Graph Completion: What is Needed to Improve Coverage?
abstract
This survey/position paper discusses ways to improve coverage of resources such as Word-Net.Rapp estimated correlations, ρ, between corpus statistics and psycholinguistic norms.ρ improves with quantity (corpus size) and quality (balance).1M words are enough for simple estimates (unigram frequencies), but at least 100M are required for pairs of words (word associations, edges).Knowledge Graph Completion (KGC) attempts to learn missing links in WN18.Unfortunately, WN18 is flawed with information leaking from train to test.More seriously, WN18 is based on SemCor (just 200k words) and dated (collected in 1960s).KGC cannot learn anything that happened since the 1960s, or associations requiring 100M words.
Kenneth Church 0001, Yuchen Bian
EMNLP (1)2
2021 Isotropy in the Contextual Embedding Space: Clusters and Manifolds
Xingyu Cai, Jiaji Huang, Yuchen Bian, Kenneth Church 0001
ICLR3
2021 Automatic Channel Pruning with Hyper-parameter Search and Dynamic Masking
abstract
Modern deep neural network models tend to be large and computationally intensive. One typical solution to this issue is model pruning. However, most current model pruning algorithms depend on hand crafted rules or need to input the pruning ratio beforehand. To overcome this problem, we propose a learning based automatic channel pruning algorithm for deep neural network, which is inspired by recent automatic machine learning (Auto ML). A two objectives' pruning problem that aims for the weights and the remaining channels for each layer is first formulated. An alternative optimization approach is then proposed to derive the channel numbers and weights simultaneously. In the process of pruning, we utilize a searchable hyper-parameter, remaining ratio, to denote the number of channels in each convolution layer, and then a dynamic masking process is proposed to describe the corresponding channel evolution. To adjust the trade-off between accuracy of a model and the pruning ratio of floating point operations, a new loss function is further introduced. Extensive experimental results on benchmark datasets demonstrate that our scheme achieves competitive results for neural network pruning.
Baopu Li, Yanwen Fan, Zhihong Pan 0001, Yuchen Bian
ACM Multimedia4
2021 On Attention Redundancy: A Comprehensive Study
abstract
Yuchen Bian, Jiaji Huang, Xingyu Cai, Jiahong Yuan, Kenneth Church. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Yuchen Bian, Jiaji Huang, Xingyu Cai, Jiahong Yuan, Kenneth Church 0001
NAACL-HLT1
2020 Rethinking Local Community Detection: Query Nodes Replacement
abstract
Local community detection for a given set of query nodes attracts much research attention recently. The query nodes play essential roles in the detection effectiveness. Existing methods perform well when a query node is from the target community core region. However, they struggle with the query-bias issue and especially perform unsatisfactorily when the query nodes come from different communities or when certain query nodes are from communities overlapping region or community boundary region. To address above issues, we consider from a new angle, to replace these original “intractable” query nodes with new detection-friendly query nodes. In this paper, we propose an effective ATP (Amplified Topology Potential) algorithm to detect core nodes of the target communities w.r.t. original query nodes. For one query node, ATP first builds a query-oriented topology potential field around the query node by aggregating random walk with restart scores. Then it amplifies the topology potential value to make core nodes of target communities easily distinguished. Graph-size-independent fast approximation strategies are also proposed together with sound theoretical foundations. Extensive experiments on four real networks using ten state-of-the-art local community detection methods verify the improvement in detection effectiveness and efficiency by the replacing strategy for the tough query cases.
Yuchen Bian, Jun Huan, Dejing Dou, Xiang Zhang 0001
ICDM1
2020 Disfluencies and Fine-Tuning Pre-Trained Language Models for Detection of Alzheimer's Disease
Jiahong Yuan, Yuchen Bian, Xingyu Cai, Jiaji Huang, Kenneth Church 0001
INTERSPEECH2
2020 Local Community Detection in Multiple Networks
abstract
Local community detection aims to find a set of densely-connected nodes containing given query nodes. Most existing local community detection methods are designed for a single network. However, a single network can be noisy and incomplete. Multiple networks are more informative in real-world applications. There are multiple types of nodes and multiple types of node proximities. Complementary information from different networks helps to improve detection accuracy. In this paper, we propose a novel RWM (Random Walk in Multiple networks) model to find relevant local communities in all networks for a given query node set from one network. RWM sends a random walker in each network to obtain the local proximity w.r.t. the query nodes (i.e., node visiting probabilities).
Yuchen Bian, Yaowei Yan, Xiao Liu 0039, Jun Huan, Xiang Zhang 0001
KDD2
2020 Deep Multi-Graph Clustering via Attentive Cross-Graph Association
abstract
Multi-graph clustering aims to improve clustering accuracy by leveraging information from different domains, which has been shown to be extremely effective for achieving better clustering results than single graph based clustering algorithms. Despite the previous success, existing multi-graph clustering methods mostly use shallow models, which are incapable to capture the highly non-linear structures and the complex cluster associations in multi-graph, thus result in sub-optimal results. Inspired by the powerful representation learning capability of neural networks, in this paper, we propose an end-to-end deep learning model to simultaneously infer cluster assignments and cluster associations in multi-graph. Specifically, we use autoencoding networks to learn node embeddings. Meanwhile, we propose a minimum-entropy based clustering strategy to cluster nodes in the embedding space for each graph. We introduce two regularizers to leverage both within-graph and cross-graph dependencies. An attentive mechanism is further developed to learn cross-graph cluster associations. Through extensive experiments on a variety of datasets, we observe that our method outperforms state-of-the-art baselines by a large margin.
Jingchao Ni, Suhang Wang, Yuchen Bian, Xiong Bill Yu, Xiang Zhang 0001
WSDM4
2020 Memory-based random walk for multi-query local community detection
Yuchen Bian, Yaowei Yan, Wei Cheng 0002, Wei Wang 0010, Xiang Zhang 0001
Knowl. Inf. Syst.1
2020 Correction to: Memory-based random walk for multi-query local community detection
Yuchen Bian, Yaowei Yan, Wei Cheng 0002, Wei Wang 0010, Xiang Zhang 0001
Knowl. Inf. Syst.1
2019 Constrained Local Graph Clustering by Colored Random Walk
abstract
Detecting local graph clusters is an important problem in big graph analysis. Given seed nodes in a graph, local clustering aims at finding subgraphs around the seed nodes, which consist of nodes highly relevant to the seed nodes. However, existing local clustering methods either allow only a single seed node, or assume all seed nodes are from the same cluster, which is not true in many real applications. Moreover, the assumption that all seed nodes are in a single cluster fails to use the crucial information of relations between seed nodes. In this paper, we propose a method to take advantage of such relationship. With prior knowledge of the community membership of the seed nodes, the method labels seed nodes in the same (different) community by the same (different) color. To further use this information, we introduce a color-based random walk mechanism, where colors are propagated from the seed nodes to every node in the graph. By the interaction of identical and distinct colors, we can enclose the supervision of seed nodes into the random walk process. We also propose a heuristic strategy to speed up the algorithm by more than 2 orders of magnitude. Experimental evaluations reveal that our clustering method outperforms state-of-the-art approaches by a large margin.
Yaowei Yan, Yuchen Bian, Dongwon Lee 0001, Xiang Zhang 0001
WWW2
2019 The multi-walker chain and its application in local community detection
Yuchen Bian, Jingchao Ni, Wei Cheng 0002, Xiang Zhang 0001
Knowl. Inf. Syst.1
2018 On Multi-query Local Community Detection
abstract
Local community detection, which aims to find a target community containing a set of query nodes, has recently drawn intense research interest. The existing local community detection methods usually assume all query nodes are from the same community and only find a single target community. This is a strict requirement and does not allow much flexibility. In many real-world applications, however, we may not have any prior knowledge about the community memberships of the query nodes, and different query nodes may be from different communities. To address this limitation of the existing methods, we propose a novel memory-based random walk method, MRW, that can simultaneously identify multiple target local communities to which the query nodes belong. In MRW, each query node is associated with a random walker. Different from commonly used memoryless random walk models, MRW records the entire visiting history of each walker. The visiting histories of walkers can help unravel whether they are from the same community or not. Intuitively, walkers with similar visiting histories are more likely to be in the same community. Moreover, MRW allows walkers with similar visiting histories to reinforce each other so that they can better capture the community structure instead of being biased to the query nodes. We provide rigorous theoretical foundation for the proposed method and develop efficient algorithms to identify multiple target local communities simultaneously. Comprehensive experimental evaluations on a variety of real-world datasets demonstrate the effectiveness and efficiency of the proposed method.
Yuchen Bian, Yaowei Yan, Wei Cheng 0002, Wei Wang 0010, Xiang Zhang 0001
ICDM1
2018 Second-order random walk-based proximity measures in graph analysis: formulations and algorithms
Yubao Wu, Xiang Zhang 0001, Yuchen Bian, Zhipeng Cai 0001, Xiang Lian 0001, Xueting Liao, Fengpan Zhao
VLDB J.3
2017 Many Heads are Better than One: Local Community Detection by the Multi-walker Chain
abstract
Local community detection (or local clustering) is of fundamental importance in large network analysis. Random walk based methods have been routinely used in this task. Most existing random walk methods are based on the single-walker model. However, without any guidance, a single-walker may not be adequate to effectively capture the local cluster. In this paper, we study a multi-walker chain (MWC) model, which allows multiple walkers to explore the network. Each walker is influenced (or pulled back) by all other walkers when deciding the next steps. This helps the walkers to stay as a group and within the cluster. We introduce two measures based on the mean and standard deviation of the visiting probabilities of the walkers. These measures not only can accurately identify the local cluster, but also help detect the cluster center and boundary, which cannot be achieved by the existing single-walker methods. We provide rigorous theoretical foundation for MWC, and devise efficient algorithms to compute it. Extensive experimental results on a variety of real-world networks demonstrate that MWC outperforms the state-of-the-art local community detection methods by a large margin.
Yuchen Bian, Jingchao Ni, Wei Cheng 0002, Xiang Zhang 0001
ICDM1
2016 Remember Where You Came From: On The Second-Order Random Walk Based Proximity Measures
abstract
Measuring the proximity between different nodes is a fundamental problem in graph analysis. Random walk based proximity measures have been shown to be effective and widely used. Most existing random walk measures are based on the first-order Markov model, i.e., they assume that the next step of the random surfer only depends on the current node. However, this assumption neither holds in many real-life applications nor captures the clustering structure in the graph. To address the limitation of the existing first-order measures, in this paper, we study the second-order random walk measures, which take the previously visited node into consideration. While the existing first-order measures are built on node-to-node transition probabilities, in the second-order random walk, we need to consider the edge-to-edge transition probabilities. Using incidence matrices, we develop simple and elegant matrix representations for the second-order proximity measures. A desirable property of the developed measures is that they degenerate to their original first-order forms when the effect of the previous step is zero. We further develop Monte Carlo methods to efficiently compute the second-order measures and provide theoretical performance guarantees. Experimental results show that in a variety of applications, the second-order measures can dramatically improve the performance compared to their first-order counterparts.
Yubao Wu, Yuchen Bian, Xiang Zhang 0001
Proc. VLDB Endow.2