EDBT 2026 Demo / reviewers in the wild / expert
Xiaoran Yan
dblp:05/8149
· DBLP profile ↗
29ranked-venue papers
2as first author
23since 2021 · last 2026
0000-0003-3481-1832ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 11 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 8 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 4 since 2021Security and privacy · 3 · 3 since 2021Computer networks · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DGNet: Enhancing Parallel CTR Prediction Models via Decoupled Gated Network: Enhancing Parallel CTR Prediction Models via Decoupled Gated NetworkabstractClick-through rate (CTR) prediction is a cornerstone of modern recommender systems and online advertising platforms. Parallel CTR models utilize multiple sub-networks to capture diverse feature interactions and achieve advanced performance. However, these models face two key limitations: (1) reliance on shared, static feature embeddings limits their ability to capture distinct interaction signals across parallel sub-networks; and (2) although effective for modeling high-order interactions, the conventional layer-by-layer interaction paradigm propagates and amplifies noise cumulatively, thereby reducing model robustness. To address these challenges, we propose Decoupled Gated Network (DGNet), a novel framework that introduces two key components: Decoupled Embedding Generator (DEG) adaptively generates sub-network-specific embeddings from original representations, enhancing interaction specificity. Gated Fusion Module (GateF) dynamically extracts layer-wise complementary information from decoupled embeddings to mitigate cumulative interaction noise. DGNet and its two components are model-agnostic and can be seamlessly integrated into existing parallel CTR models. Extensive experiments show that DGNet consistently improves the performance of various parallel CTR models, yielding statistically significant performance boosting. Meanwhile, visualization and quantitative analyses provide an intuitively explanation for its reliability and performance improvements. Fangye Wang, Xiaoran Yan |
WSDM | 2 |
| 2025 | Hierarchical Action Recognition: A Contrastive Video-Language Approach with Hierarchical Interactions
Rui Zhang 0055, Shuailong Li, Junxiao Xue, Diying Yan, Xiaoran Yan |
CGI (2) | 8 |
| 2025 | Comparing and Improving Perturbation Mechanisms Under Local Differential Privacy
She Sun, Xiaoran Yan, Huiwen Wu |
Inscrypt (3) | 3 |
| 2025 | Alignment-Uniformity Aware Feature Representation Learning for CTR Prediction
Fangye Wang, Xiaoran Yan |
DASFAA (5) | 2 |
| 2025 | High-Resolution Face Reconstruction via Gaussian Splatting: Seeing Through Deblur
Dapeng Zhao, Xiaoran Yan |
PRCV (10) | 2 |
| 2025 | Hier-FUN: Hierarchical Federated Learning and Unlearning in Heterogeneous Edge ComputingabstractFederated learning (FL) has emerged as a pivotal paradigm for distributed model training in edge computing (EC), enabling cooperation among numerous Internet of Things devices while safeguarding their data privacy. Despite its successes in machine learning, concerns regarding data security and model fidelity necessitate the efficient unlearning of target device, i.e., federated unlearning (FUN). However, due to resource constraints, device heterogeneity, and non-independent and identically distributed (Non-IID) data, securely eliminating a device’s impact without retraining the model from scratch presents a complex challenge. In response to these challenges, we propose a hierarchical FUN framework, called Hier-FUN. Hier-FUN organizes edge devices into K clusters, each managed by a head device responsible for aggregating local models within the cluster. To expedite both the learning and unlearning processes of Hier-FUN, we design a heuristic algorithm to determine an appropriate value for K based on devices’ data distributions and available resources. In addition, Hier-FUN denies the communication between the server and cluster heads during training, which can constrain the influence sphere of target device and accelerate the unlearning process. We conduct extensive experiments using real-world datasets, and the experimental results illustrate that Hier-FUN can improve test accuracy by 3.19% during the learning phase and achieve a$6.8\times $speedup during unlearning compared with the baseline methods. Zhen-guo Ma, Huaqing Tu, Pengli Ji, Xiaoran Yan, Hongli Xu 0001, Zhiyuan Wang 0002, Suo Chen |
IEEE Internet Things J. | 5 |
| 2025 | VFGCN: A Vertical Federated Learning Framework With Privacy Preserving for Graph Convolutional NetworkabstractDue to the robust representational capabilities of graph data, employing graph neural networks for its processing has demonstrated superior performance over conventional deep learning algorithms. Graph data encompasses abundant features and structural information; however, its large-scale collection is often challenging in practice. This difficulty arises because data predominantly exists in isolated compartments, making it arduous to harmonize information across various organizations or to enable multiple organizations to collaborate effectively while safeguarding local data privacy. In light of an extreme data distribution scenario, where each client possesses distinct nodes with partially overlapping segments yet divergent data features, we introduce a dual-cloud server architecture. This framework encompasses the design of four secure subprotocols: ReEnc (secure re-encryption), SecPSI (secure outsourcing of PSI), SecWeight (secure weight calculation), and SecAgg (secure aggregation). Together, these components facilitate a vertical federated learning framework for graph convolutional networks, ensuring privacy preservation. We provide a security proof for the entire system and extensive evaluation on three benchmark datasets (Cora, Citeseer, and Pubmed) illustrates that our Vertical Federated Graph Convolutional Network (VFGCN) surpasses existing privacy-preserving methodologies. Qingming Li, Ximeng Liu, Xiaoran Yan, Qingkuan Dong, Huiwen Wu, Xiangjie Kong 0001 |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2025 | Federated Graph Anomaly Detection via Contrastive Self-Supervised LearningabstractAttribute graph anomaly detection aims to identify nodes that significantly deviate from the majority of normal nodes, and has received increasing attention due to the ubiquity and complexity of graph-structured data in various real-world scenarios. However, current mainstream anomaly detection methods are primarily designed for centralized settings, which may pose privacy leakage risks in certain sensitive situations. Although federated graph learning offers a promising solution by enabling collaborative model training in distributed systems while preserving data privacy, a practical challenge arises as each client typically possesses a limited amount of graph data. Consequently, naively applying federated graph learning directly to anomaly detection tasks in distributed environments may lead to suboptimal performance results. We propose a federated graph anomaly detection framework via contrastive self-supervised learning (CSSL) [federated CSSL anomaly detection framework (FedCAD)] to address these challenges. FedCAD updates anomaly node information between clients via federated learning (FL) interactions. First, FedCAD uses pseudo-label discovery to determine the anomaly node of the client preliminarily. Second, FedCAD employs a local anomaly neighbor embedding aggregation strategy. This strategy enables the current client to aggregate the neighbor embeddings of anomaly nodes from other clients, thereby amplifying the distinction between anomaly nodes and their neighbor nodes. Doing so effectively sharpens the contrast between positive and negative instance pairs within contrastive learning, thus enhancing the efficacy and precision of anomaly detection through such a learning paradigm. Finally, the efficiency of FedCAD is demonstrated by experimental results on four real graph datasets. Xiangjie Kong 0001, Hui Wang 0097, Mingliang Hou, Xin Chen 0054, Xiaoran Yan, Sajal K. Das 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2024 | Point-Correlate Adversarial Transformer for Unsupervised Multivariate Time Series Anomaly DetectionabstractMultivariate time series anomaly detection plays a crucial role in industrial production. However, the inherent complexity and randomness of time series pose significant challenges. Furthermore, existing detection methods struggle to provide reliable explanations for outliers. To address these issues, this paper presents an unsupervised multivariate time series anomaly detection model named Point-Correlate Adversarial Transformer (PCAT). In this work, we leverage Transformer networks to capture the underlying correlations between different points in a time series and reconstruct the original sequence. By analyzing the correlation differences and reconstruction errors, we identify anomalies at the point level. Our model incorporates an adversarial structure, enabling unsupervised learning and enhancing the learning capability and robustness of the detection network. Experimental evaluations on four real-world datasets demonstrate the superiority of our approach over other state-of-the-art models in terms of detection delay and accuracy. Xiangjie Kong 0001, Guojiang Shen, Xiaoran Yan, Mario Collotta |
CSCWD | 4 |
| 2024 | AdaFL: Adaptive Client Selection and Dynamic Contribution Evaluation for Efficient Federated LearningabstractFederated learning is a collaborative machine learning framework where multiple clients jointly train a global model. To mitigate communication overhead, it is common to select a subset of clients for participation in each training round. However, existing client selection strategies often rely on a fixed number of clients throughout all rounds, which may not be the optimal choice for balancing training efficiency and model performance. Moreover, these approaches typically evaluate clients solely based on their performances in one single round, neglecting the effects of historical records and potentially introducing randomness into the global model. In our work, we introduce AdaFL, a novel approach to client selection and contribution evaluation for efficient federated learning. AdaFL dynamically adjusts the number of clients to be selected using a piecewise function. It initiates with a small selection size to reduce communication overhead and progressively increases it to enhance model generalization. Furthermore, AdaFL evaluates clients’ contributions by combining their performance metrics from both current and historical rounds through a weighted average function, with a weight parameter fine-tuning the trade-off between current and historical data. Experimental results show that the proposed AdaFL outperforms prior works in terms of improving test accuracy and reducing training runtime. Qingming Li, Xiaoran Yan |
ICASSP | 4 |
| 2024 | Position-Aware Active Learning for Multi-Modal Entity AlignmentabstractMulti-Modal Entity Alignment (MMEA) aims to identify equivalent entities across different knowledge graphs by utilizing auxiliary modalities such as images. While MMEA has made significant progress, prevailing methods still heavily rely on abundant annotated entity pairs. Active learning seeks to alleviate the labeling burden or enhance model efficiency within fixed labeling capacity through careful sample selection. However, active learning for entity alignment in multimodal scenarios remains unexplored. In our view, it is crucial that data selected from different modalities should complement each other without redundancy or overlap; otherwise, the obtained data may prove a waste of labeling budgets. To achieve this goal, we propose a novel acquisition function leveraging Graph Neural Networks’ (GNNs) capability to aggregate information over multiple hops, prioritizing data distant from other modalities’ selections. Moreover, existing approaches employ data augmentation by selecting entity pairs whose inter-entity similarities of other modalities exceed a predefined threshold, but this augmentation strategy inadequately capitalizes on the available similarity information among entities. We can further enhance performance by integrating similarity matrices from different modalities. Consequently, our method achieves considerable improvements over existing active learning methods for entity alignment, as demonstrated by the experiments. Baogui Xu, Yafei Lu, Bing Su 0001, Xiaoran Yan |
ICASSP | 4 |
| 2024 | Video-Language Graph Convolutional Network for Human Action RecognitionabstractTransferring visual language models (VLMs) from the image domain to the video domain has recently yielded great success on human action recognition tasks. However, standard recognition paradigms overlook fine-grained action parsing knowledge that could enhance the recognition accuracy. In this paper, we propose a novel method that leverages both coarse-grained and fine-grained knowledge to recognize human actions in videos. Our method consists of a video-language graph convolutional network that integrates and fuses multi-modal knowledge in a progressive manner. We evaluate our method on the Kinetics-TPS, a large-scale action parsing dataset, and demonstrate that it outperforms the state-of-the-art methods by a significant margin. Moreover, our method achieves better results with less training data and competitive computational cost than the existing methods, showing the effectiveness and efficiency of using fine-grained knowledge for human video action recognition. Rui Zhang 0055, Xiaoran Yan |
ICASSP | 2 |
| 2024 | Enhancing Human Action Recognition with Fine-grained Body Movement AttentionabstractIn the field of vision-language models (VLMs), human action recognition models, while effective, always rely on large pre-trained models or high-resolution inputs, leading to computational challenges. To address this, we propose a novel VLM approach with fine-grained attention to body movements. Unlike methods relying on coarse video-text matching, we guide the model to infer actions from fine-grained body part movements using two techniques: fine-tuning pre-trained encoders at the fine-grained level and matching labels from language and vision perspectives at the coarse-grained level. Experiments show our model excels in fully-supervised, few-shot, and zero-shot scenarios with just 8 random frames and a ViT-B/32 backbone. It outperforms most ViT-L/14 based models, demonstrating effectiveness while saving computational resources. The largest Top-1 accuracy improvement over second-best approaches is 6.8%. Rui Zhang 0055, Junxiao Xue, Pavel Smirnov 0005, Xiaoran Yan |
ICME | 7 |
| 2024 | Flexible Graph Neural Diffusion with Latent Class Representation LearningabstractIn existing graph data, the connection relationships often exhibit uniform weights, leading to the model aggregating neighboring nodes with equal weights across various connection types. However, this uniform aggregation of diverse information diminishes the discriminability of node representations, contributing significantly to the over-smoothing issue in models. In this paper, we propose the Flexible Graph Neural Diffusion (FGND) model, incorporating latent class representation to address the misalignment between graph topology and node features. In particular, we combine latent class representation learning with the inherent graph topology to reconstruct the diffusion matrix during the graph diffusion process. We introduce the sim metric to quantify the degree of mismatch between graph topology and node features. By flexibly adjusting the dependency level on node features through the hyperparameter, we accommodate diverse adjacency relationships. The effective filtering of noise in the topology also allows the model to capture higher order information, significantly alleviating the over-smoothing problem. Meanwhile, we model the graphical diffusion process as a set of differential equations and employ advanced partial differential equation tools to obtain more accurate solutions. Empirical evaluations on five benchmarks reveal that our FGND model outperforms existing popular GNN methods in terms of both overall performance and stability under data perturbations. Meanwhile, our model exhibits superior performance in comparison to models tailored for heterogeneous graphs and those designed to address oversmoothing issues. Liangtian Wan, Huijin Han, Lu Sun 0004, Zixun Zhang, Zhaolong Ning, Xiaoran Yan, Feng Xia 0001 |
KDD | 6 |
| 2024 | FedSGProx: Mitigating Data Heterogeneity and Isolated Nodes in Graph Federated LearningabstractGraphs capture complex node interactions and are a fundamental tool for machine learning. Graph Federated Learning (GFL) is a method that allows multiple clients to collaboratively train a global graph neural network using a federated learning framework. This approach leverages the value of distributed graph data while maintaining data privacy. Existing approaches optimize graph neural networks within the common FedAvg paradigm, but they face two problems. The first problem is data heterogeneity, which leads to variations in label distributions and subgraph structures across clients. The second problem involves isolated nodes that possess limited or no local connections. The two problems seriously degrade the performance of GFL. To address these issues, we introduce FedSGProx, a novel GFL approach. FedSGProx combines longterm and short-term constraints to mitigate local biases due to data heterogeneity. Moreover, we design a novel sampling strategy to limit the involvement of isolated nodes in local training, thereby reducing their negative impacts on local models. Empirical results show that FedSGProx achieves higher classification accuracy than existing methods, and its performance is very close to that in centralized training. Xutao Meng, Qingming Li, Xiaoran Yan |
TrustCom | 5 |
| 2024 | FVFL: A Flexible and Verifiable Privacy-Preserving Federated Learning SchemeabstractWith the development of deep learning, people are more and more concerned about the security of data. Federated learning can solve the problem of data island, but it also brings more serious data privacy problems. Furthermore, in the process of multisource data collaboration, the efficiency of the whole federated learning system is usually not high. In this article, we introduce a scheme named FVFL, which ensure the local data security and resistance to collusive attacks, more importantly it can well support client flexible participate federated learning. We adopt Paillier encryption and secret sharing to guarantee client’s data security and resistance to collusive attacks. Moreover, our encryption mechanism allows client to participate in federated learning flexibly, and the correctness of the encryption algorithm is not affected by client’s drop out. The super-increasing sequence is introduced to reduce the communication overhead of the whole system, the simulation result shows that the result is significant; the Lagrange interpolation polynomial and secret Sharing is introduced to implement verification mechanism, to prevent malicious forgery of aggregation results in the cloud. The verification mechanism ensures the clients to obtain real and reliable aggregation results in the cloud. Moreover, our verification mechanism allows client to participate in federated learning flexibly, and the correctness of the verification algorithm is not affected by client’s drop out. And the experimental results show that FVFL has high accuracy and efficiency. Qingming Li, Xiaoran Yan, Ximeng Liu, Yuncheng Wu |
IEEE Internet Things J. | 4 |
| 2024 | Z-Laplacian Matrix Factorization: Network Embedding With Interpretable Graph SignalsabstractNetwork embedding aims to represent nodes with low dimensional vectors while preserving structural information. It has been recently shown that many popular network embedding methods can be transformed into matrix factorization problems. In this paper, we propose the unifying framework “Z-NetMF,” which generalizes random walk samplers to Z-Laplacian graph filters, leading to embedding algorithms with interpretable parameters. In particular, by controlling biases in the time domain, we propose the Z-NetMF-t algorithm, making it possible to scale contributions of random walks of different length. Inspired by node2vec, we design the Z-NetMF-g algorithm, capturing the random walk biases in the graph domain. Moreover, we evaluate the effect of the bias parameters based on node classification and link prediction tasks. The results show that our algorithms, especially the combined model Z-NetMF-gt with biases in both domains, outperform the state-of-art methods while providing interpretable insights at the same time. Finally, we discuss future directions of the Z-NetMF framework. Liangtian Wan, Zhengqiang Fu, Yi Ling, Lu Sun 0004, Feng Xia 0001, Xiaoran Yan, Charu C. Aggarwal |
IEEE Trans. Knowl. Data Eng. | 8 |
| 2023 | Self-Supervised Teaching and Learning of Representations on GraphsabstractRecent years have witnessed significant advances in graph contrastive learning (GCL), while most GCL models use graph neural networks as encoders based on supervised learning. In this work, we propose a novel graph learning model called GraphTL, which explores self-supervised teaching and learning of representations on graphs. One critical objective of GCL is to retain original graph information. For this purpose, we design an encoder based on the idea of unsupervised dimensionality reduction of locally linear embedding (LLE). Specifically, we map one iteration of the LLE to one layer of the network. To guide the encoder to better retain the original graph information, we propose an unbalanced contrastive model consisting of two views, which are the learning view and the teaching view, respectively. Furthermore, we consider the nodes that are identical in muti-views as positive node pairs, and design the node similarity scorer so that the model can select positive samples of a target node. Extensive experiments have been conducted over multiple datasets to evaluate the performance of GraphTL in comparison with baseline models. Results demonstrate that GraphTL can reduce distances between similar nodes while preserving network topological and feature information, yielding better performance in node classification. Liangtian Wan, Zhenqiang Fu, Lu Sun 0004, Xianpeng Wang 0001, Gang Xu 0002, Xiaoran Yan, Feng Xia 0001 |
WWW | 6 |
| 2023 | Unifying and Improving Graph Convolutional Neural Networks with Wavelet Denoising FiltersabstractGraph convolutional neural network (GCN) is a powerful deep learning framework for network data. However, variants of graph neural architectures can lead to drastically different performance on different tasks. Model comparison calls for a unifying framework with interpretability and principled experimental procedures. Based on the theories from graph signal processing (GSP), we show that GCN’s capability is fundamentally limited by the uncertainty principle, and wavelets provide a controllable trade-off between local and global information. We adapt wavelet denoising filters to the graph domain, unifying popular variants of GCN under a common interpretable mathematical framework. Furthermore, we propose WaveThresh and WaveShrink which are novel GCN models based on proven denoising filters from the signal processing literature. Empirically, we evaluate our models and other popular GCNs under a more principled procedure and analyze how trade-offs between local and global graph signals can lead to better performance in different datasets. Liangtian Wan, Huijin Han, Xiaoran Yan, Lu Sun 0004, Zhaolong Ning, Feng Xia 0001 |
WWW | 4 |
| 2023 | DPVisCreator: Incorporating Pattern Constraints to Privacy-preserving Visualizations via Differential PrivacyabstractData privacy is an essential issue in publishing data visualizations. However, it is challenging to represent multiple data patterns in privacy-preserving visualizations. The prior approaches target specific chart types or perform an anonymization model uniformly without considering the importance of data patterns in visualizations. In this paper, we propose a visual analytics approach that facilitates data custodians to generate multiple private charts while maintaining user-preferred patterns. To this end, we introduce pattern constraints to model users' preferences over data patterns in the dataset and incorporate them into the proposed Bayesian network-based Differential Privacy (DP) model PriVis. A prototype system, DPVisCreator, is developed to assist data custodians in implementing our approach. The effectiveness of our approach is demonstrated with quantitative evaluation of pattern utility under the different levels of privacy protection, case studies, and semi-structured expert interviews. Jiehui Zhou, Xumeng Wang, Jason K. Wong, Huanliang Wang, Xiaoran Yan, Haozhe Feng, Huamin Qu, Haochao Ying, Wei Chen 0001 |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2022 | Deep Reinforcement Learning-Based Energy-Efficient Edge Computing for Internet of VehiclesabstractMobile network operators (MNOs) allocate computing and caching resources for mobile users by deploying a central control system. Existing studies mainly use programming and heuristic methods to solve the resource allocation problem, which ignores the energy cost problem that is really significant to the MNO. To solve this problem, in this article, we design a joint computing and caching framework by integrating deep deterministic policy gradient (DDPG) algorithm. Especially, we focus on the Internet of Vehicles scenario, which needs the support of mobile network provided by MNO. We first formulate an optimization problem to minimize MNO’s energy cost by considering the computation and caching energy costs jointly. Then, we turn the formulated problem into a reinforcement learning problem and utilize DDPG methods to solve this problem. The final simulation result shows that our solution can reduce energy costs by more than 15%, while ensuring the tasks can be completed on time. Xiangjie Kong 0001, Gaohui Duan, Mingliang Hou, Guojiang Shen, Hui Wang 0097, Xiaoran Yan, Mario Collotta |
IEEE Trans. Ind. Informatics | 6 |
| 2021 | CADRE: A Cloud-Based Data Service for Big Bibliographic DataabstractLarge bibliographic data sets hold the promise of revolutionizing the scientific enterprise when combined with state-of-the-science computational capabilities. Providing high-quality data services for large network datasets such as the Microsoft Academic Graph, which contains more than two billion citation links, poses significant difficulties for universities. Data systems based on the property graph model are capable of delivering efficient graph query services for large networks. However, real-life queries often combine multiple types of data models. To satisfy the needs of different user groups, we developed and deployed a cloud-based data system consisting of scalable graph and text-indexed query engines. For non-expert users, the property graph model also presents a technological barrier. To alleviate the steep learning curve, we designed an intuitive graphical user interface for query-building. For advanced users, a scalable notebook service in our platform provides a more flexible computing environments where the query results can be further analyzed. These systems form the data-backbone of the Collaborative Archive and Data Research Environment (CADRE), which provides efficient and high-quality bibliographic data services to eleven large public universities in North America. Xiaoran Yan, Guangchen Ruan, Dimitar Nikolov, Matthew Hutchinson, Chathuri Peli Kankanamalage, Benjamin Serrette, James R. McCombs, Alan Walsh, Esen Tuna, Valentin Pentchev |
CIKM | 1 |
| 2021 | Detecting Outlier Patterns With Query-Based Artificially Generated Searching ConditionsabstractIn the age of social computing, finding interesting network patterns or motifs is significant and critical for various areas, such as decision intelligence, intrusion detection, medical diagnosis, social network analysis, fake news identification, and national security. However, subgraph matching remains a computationally challenging problem, let alone identifying special motifs among them. This is especially the case in large heterogeneous real-world networks. In this article, we propose an efficient solution for discovering and ranking human behavior patterns based on network motifs by exploring a user's query in an intelligent way. Our method takes advantage of the semantics provided by a user's query, which in turn provides the mathematical constraint that is crucial for faster detection. We propose an approach to generate query conditions based on the user's query. In particular, we use meta paths between the nodes to define target patterns as well as their similarities, leading to efficient motif discovery and ranking at the same time. The proposed method is examined in a real-world academic network using different similarity measures between the nodes. The experiment result demonstrates that our method can identify interesting motifs and is robust to the choice of similarity measures. Shuo Yu 0001, Feng Xia 0001, Tao Tang 0007, Xiaoran Yan, Ivan Lee 0001 |
IEEE Trans. Comput. Soc. Syst. | 5 |
| 2020 | TBI2Flow: Travel behavioral inertia based long-term taxi passenger flow prediction
Xiangjie Kong 0001, Feng Xia 0001, Zhenhuan Fu, Xiaoran Yan, Amr Tolba, Zafer Al-Makhadmeh |
World Wide Web | 4 |
| 2019 | A spectrum of routing strategies for brain networksabstractCommunication of signals among nodes in a complex network poses fundamental problems of efficiency and cost. Routing of messages along shortest paths requires global information about the topology, while spreading by diffusion, which operates according to local topological features, is informationally "cheap" but inefficient. We introduce a stochastic model for network communication that combines local and global information about the network topology to generate biased random walks on the network. The model generates a continuous spectrum of dynamics that converge onto shortest-path and random-walk (diffusion) communication processes at the limiting extremes. We implement the model on two cohorts of human connectome networks and investigate the effects of varying the global information bias on the network's communication cost. We identify routing strategies that approach a (highly efficient) shortest-path communication process with a relatively small global information bias on the system's dynamics. Moreover, we show that the cost of routing messages from and to hub nodes varies as a function of the global information bias driving the system's dynamics. Finally, we implement the model to identify individual subject differences from a communication dynamics point of view. The present framework departs from the classical shortest paths vs. diffusion dichotomy, unifying both models under a single family of dynamical processes that differ by the extent to which global information about the network topology influences the routing patterns of neural signals traversing the network. Andrea Avena-Koenigsberger, Xiaoran Yan, Artemy Kolchinsky, Martijn P. van den Heuvel, Patric Hagmann, Olaf Sporns |
PLoS Comput. Biol. | 2 |
| 2016 | Bayesian model selection of stochastic block modelsabstractA central problem in analyzing networks is partitioning them into modules or communities. One of the best tools for this is the stochastic block model, which clusters vertices into blocks with statistically homogeneous pattern of links. Despite its flexibility and popularity, there has been a lack of principled statistical model selection criteria for the stochastic block model. Here we propose a Bayesian framework for choosing the number of blocks as well as comparing it to the more elaborate degree-corrected block models, ultimately leading to a universal model selection framework capable of comparing multiple modeling combinations. We will also investigate its theoretic connection to the minimum description length principle. Xiaoran Yan |
ASONAM | 1 |
| 2014 | The interplay between dynamics and networks: centrality, communities, and cheeger inequalityabstractWe study the interplay between a dynamic process and the structure of the network on which it is defined. Specifically, we examine the impact of this interaction on the quality-measure of network clusters and node centrality. This enables us to effectively identify network communities and important nodes participating in the dynamics. As the first step towards this objective, we introduce an umbrella framework for defining and characterizing an ensemble of dynamic processes on a network. This framework generalizes the traditional Laplacian framework to continuous-time biased random walks and also allows us to model some epidemic processes over a network. For each dynamic process in our framework, we can define a function that measures the quality of every subset of nodes as a potential cluster (or community) with respect to this process on a given network. This subset-quality function generalizes the traditional conductance measure for graph partitioning. We partially justify our choice of the quality function by showing that the classic Cheeger's inequality, which relates the conductance of the best cluster in a network with a spectral quantity of its Laplacian matrix, can be extended from the Laplacian-conductance setting to this more general setting. Rumi Ghosh, Shang-Hua Teng, Kristina Lerman, Xiaoran Yan |
KDD | 4 |
| 2013 | Scalable text and link analysis with mixed-topic link modelsabstractMany data sets contain rich information about objects, as well as pairwise relations between them. For instance, in networks of websites, scientific papers, and other documents, each node has content consisting of a collection of words, as well as hyperlinks or citations to other nodes. In order to perform inference on such data sets, and make predictions and recommendations, it is useful to have models that are able to capture the processes which generate the text at each node and the links between them. In this paper, we combine classic ideas in topic modeling with a variant of the mixed-membership block model recently developed in the statistical physics community. The resulting model has the advantage that its parameters, including the mixture of topics of each document and the resulting overlapping communities, can be inferred with a simple and scalable expectation-maximization algorithm. We test our model on three data sets, performing unsupervised topic classification and link prediction. For both tasks, our model outperforms several existing state-of-the-art methods, achieving higher accuracy with significantly less computation, analyzing a data set with 1.3 million words and 44 thousand links in a few minutes. Yaojia Zhu, Xiaoran Yan, Lise Getoor, Cristopher Moore |
KDD | 2 |
| 2011 | Active learning for node classification in assortative and disassortative networksabstractIn many real-world networks, nodes have class labels or variables that affect the network's topology. If the topology of the network is known but the labels of the nodes are hidden, we would like to select a small subset of nodes such that, if we knew their labels, we could accurately predict the labels of all the other nodes. We develop an active learning algorithm for this problem which uses information-theoretic techniques to choose which nodes to explore. We test our algorithm on networks from three different domains: a social network, a network of English words that appear adjacently in a novel, and a marine food web. Our algorithm makes no initial assumptions about how the groups connect, and performs well even when faced with quite general types of network structure. In particular, we do not assume that nodes of the same class are more likely to be connected to each other - only that they connect to the rest of the network in similar ways. Cristopher Moore, Xiaoran Yan, Yaojia Zhu, Jean-Baptiste Rouquier, Terran Lane |
KDD | 2 |