Qian Yan 0001

dblp:120/4223-1 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
6since 2021 · last 2026
0000-0003-4715-6830ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 7 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Multi-task Inference of Diffusion Networks
abstract
Inferring the underlying structures of diffusion networks based on observed diffusion results is a fundamental problem in network analysis. Traditional approaches typically address this problem by inferring each diffusion network in isolation, relying on the assumption that sufficient observation data is available for each individual inference task. However, in many real-world scenarios, it is common to observe diffusion processes occur across multiple networks with similar structures, while the amount of observable data collected on each network is often limited. In this work, we study how to infer multiple similar diffusion networks jointly with limited observation data for each network. To this end, we propose a novel iterative strategy which in turn updates the inference results for all diffusion networks by exploiting the similarity between the networks, and theoretically guarantee the monotonicity and convergence of the iterative process. Extensive experiments on both synthetic and real-world networks demonstrate that our method not only achieves superior inference accuracy compared to existing techniques, but also maintains high computational efficiency.
Ting Gan, Kudereti Kuerban, Qian Yan 0001, Zhigao Zheng 0001, Hao Huang 0001
WWW3
2025 Triangle Counting Over Large-Scale Directed Graphs
abstract
Triangle counting calculates the number of triangular structures in a graph. It is a fundamental basis for many graph algorithms such as clustering coefficient, community detection, and link prediction. When a graph is distributed or is too large to fit into one physical machine, running triangle counting across distributed machines becomes necessary. However, existing solutions are mostly designed for undirected graphs where triangles are symmetric. They cannot work for directed graphs, in which there are different types of triangles showing different local structures. This paper studies triangle counting over large-scale directed graphs. We propose a distributed triangle counting algorithm, called T-count, for directed graphs. T-count determines the edge type using a duplicating method, avoids redundant computation by reducing the size of neighboring vertices, and infers the triangle type with a lookup table. We theoretically prove the correctness guarantee of T-count and implement it over GraphX. Extensive evaluations show that T-count can efficiently handle large-scale directed graphs and benefit downstream graph analytics in real-world industrial applications such as fraud detection.
Zhigao Zheng 0001, Qian Yan 0001, Kudereti Kuerban, Ting Gan, Hao Huang 0001
HPCC3
2025 Edge-based graph neighbor filtering network for recommendation
Qian Yan 0001, Yunbo Tang
Appl. Intell.1
2025 Diffusion pattern mining
Qian Yan 0001, Yulan Yang, Ting Gan, Hao Huang 0001
Knowl. Inf. Syst.1
2024 Learning Diffusions under Uncertainty
abstract
To infer a diffusion network based on observations from historical diffusion processes, existing approaches assume that observation data contain exact occurrence time of each node infection, or at least the eventual infection statuses of nodes in each diffusion process. They determine potential influence relationships between nodes by identifying frequent sequences, or statistical correlations, among node infections. In some real-world settings, such as the spread of epidemics, tracing exact infection times is often infeasible due to a high cost; even obtaining precise infection statuses of nodes is a challenging task, since observable symptoms such as headache only partially reveal a node’s true status. In this work, we investigate how to effectively infer a diffusion network from observation data with uncertainty. Provided with only probabilistic information about node infection statuses, we formulate the problem of diffusion network inference as a constrained nonlinear regression w.r.t. the probabilistic data. An alternating maximization method is designed to solve this regression problem iteratively, and the improvement of solution quality in each iteration can be theoretically guaranteed. Empirical studies are conducted on both synthetic and real-world networks, and the results verify the effectiveness and efficiency of our approach.
Hao Huang 0001, Qian Yan 0001, Keqi Han, Ting Gan, Jiawei Jiang 0001, Quanqing Xu, Chuanhui Yang
AAAI2
2021 Statistical Inference of Diffusion Networks
abstract
To infer structures in diffusion networks, existing approaches mostly need to know not only the final infection statuses of network nodes, but also the exact times when infections occur. In contrast, in many real-world settings, such as disease propagation, monitoring exact infection times is often infeasible due to a high cost. We investigate the problem of how to learn diffusion network structures based on only the final infection statuses of nodes. Instead of utilizing sequences of timestamps to determine potential parent-child influence relationships between nodes, we propose to find influence relationships with high statistical significance. To this end, we design a probabilistic generative model of the final infection statuses to quantitatively measure the likelihood of potential structures of the objective diffusion network, taking into account network complexity. Based on this model, we can infer an appropriate number of most probable parent nodes for each node in the network. Furthermore, to reduce redundant inference computations, we are able to preclude insignificant candidate parent nodes from being considered during inferencing, if their infections have little correlation with the infections of the corresponding child nodes. Extensive experiments on both synthetic and real-world networks offer evidence that the proposed approach is effective and efficient.
Hao Huang 0001, Qian Yan 0001, Lu Chen 0001, Yunjun Gao, Christian S. Jensen
IEEE Trans. Knowl. Data Eng.2
2020 LERI: Local Exploration for Rare-Category Identification
abstract
To identify the data examples of rare categories that form small compact clusters in large data sets, existing approaches mostly require enough labeled data examples as a training set to learn a classifier, assuming that the rare-category clusters are spherical or nearly spherical. Nonetheless, a large enough training set is usually difficult to obtain in practice, and rare categories in many real-world applications often form small compact clusters with arbitrary shapes. In this paper, we investigate how to identify all data examples of a rare category with an arbitrary shape based on only one seed (i.e., a labeled rare-category data example). Instead of finding a compact and spherical local region around the seed, we locally explore the data set from the seed by continuously searching and visiting the k-nearest neighbors of each newly visited data example. The local exploration connects the data examples in the objective rare category by the relationship of k-nearest neighbors, and meanwhile, suspected external data examples are filtered out if they are not close enough to any visited data example. Experimental results on both synthetic and real-world data sets are conducted, and the results verify the effectiveness and efficiency of our approach.
Hao Huang 0001, Qian Yan 0001, Wei Lu 0015, Huaizhong Lin, Yunjun Gao, Lei Chen 0002
IEEE Trans. Knowl. Data Eng.2
2019 Learning Diffusions without Timestamps
abstract
To learn the underlying parent-child influence relationships between nodes in a diffusion network, most existing approaches require timestamps that pinpoint the exact time when node infections occur in historical diffusion processes. In many real-world diffusion processes like the spread of epidemics, monitoring such infection temporal information is often expensive and difficult. In this work, we study how to carry out diffusion network inference without infection timestamps, using only the final infection statuses of nodes in each historical diffusion process, which are more readily accessible in practice. Our main result is a probabilistic model that can find for each node an appropriate number of most probable parent nodes, who are most likely to have generated the historical infection results of the node. Extensive experiments on both synthetic and real-world networks are conducted, and the results verify the effectiveness and efficiency of our approach.
Hao Huang 0001, Qian Yan 0001, Ting Gan, Di Niu 0002, Wei Lu 0015, Yunjun Gao
AAAI2
2017 Group-Level Influence Maximization with Budget Constraint
Qian Yan 0001, Hao Huang 0001, Yunjun Gao, Wei Lu 0015, Qinming He
DASFAA (1)1
2017 False data separation for data security in smart grids
Hao Huang 0001, Qian Yan 0001, Wei Lu 0015, Zhenguang Liu, Zongpeng Li
Knowl. Inf. Syst.2
2016 Modeling for Noisy Labels of Crowd Workers
Qian Yan 0001, Hao Huang 0001, Yunjun Gao, Chen Ying, Qingyang Hu, Tieyun Qian, Qinming He
APWeb (2)1