VLDB 2026 Research / reviewers in the wild / expert
Tao Qin 0002
dblp:14/6841-2
· DBLP profile ↗
53ranked-venue papers
9as first author
26since 2021 · last 2026
0000-0003-4874-2567ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 17 · 7 first-author · 1 since 2021Artificial intelligence and machine learning · 14 · 1 first-author · 12 since 2021Databases, data management, data science and information retrieval · 10 · 8 since 2021Security and privacy · 5 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 since 2021Systems, architecture and hardware · 1Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RTCM: A Distributed Snapshot-Based Framework for Real-Time Co-Movement Mining
Chenxu Wang 0001, Jiaxing Wei, Tianyi Li 0005, Hongzhen Xiang, Junzhou Zhao, Pinghui Wang, Tao Qin 0002, Yushuai Li, Christian S. Jensen |
EDBT | 7 |
| 2026 | SGA: Self-boosting Attributed Graph Alignment via Neighborhood Consistency-based Edge EnhancementabstractGraph alignment, the task of identifying corresponding nodes across different graphs, is crucial for applications ranging from social network analysis to bioinformatics. Although most existing methods leverage graph neural networks (GNNs) to learn node embeddings for attributed graphs and match them based on node similarity, they often rely on objectives designed for node classification or link prediction. These approaches preserve node proximity within individual graphs but fail to capture cross-graph correspondence knowledge, leading to suboptimal alignment performance. Chenxu Wang 0001, Wencong Lin, Pinghui Wang, Tao Qin 0002, Wei Wang 0012, Xiaohong Guan |
KDD (1) | 4 |
| 2026 | Knowledge-Variational Contrastive Learning for RecommendationabstractRecommender systems are effective tools to alleviate the challenges posed by information overload, but data sparsity has greatly affected their performance. Knowledge Graphs (KGs) and self-supervised learning are used to alleviate the data sparsity problem. However, existing KG-enhanced self-supervised learning recommendation methods have the following limitations: (1) Generality : existing CL-based recommendation methods strongly rely on manually designed data augmentation strategies, leading to poor generality of the CL-based models. (2) Robustness : KG usually contains lots of task-irrelevant entities, and the user–item interactions constructed from implicit feedback are usually noisy. The noisy data will generate intrusive supervised and self-supervised signals and will degrade recommendation performance. To address these limitations, we propose a novel KG-enhanced self-supervised learning recommendation method, named Knowledge-Variational Contrastive Learning for Recommendation (KVCL) . Specifically, we first design an adaptively denoising mechanism to identify and prune the noisy data in the KG and user–item interaction bipartite graph. Then, we learn a normal distribution for each node by the variational auto-encoder, and sample multiple times from the learned distribution to obtain different contrastive views. Extensive experiments based on three public datasets show that KVCL achieves improved performance over state-of-the-art methods, notably with 3.13% performance gain over state-of-the-art methods on Recall@20 and NDCG@20. Furthermore, evaluations including ablation studies and detailed analyses of multi-scenarios, computational efficiency, complexity, and denoising interpretability further underscore its scalability and practical applicability. Tao Qin 0002, Pinghui Wang, Kuiyu Zhu |
ACM Trans. Knowl. Discov. Data | 2 |
| 2025 | Attention-based Conditional Random Field for Financial Fraud DetectionabstractFinancial fraud detection is critical for market transparency and regulatory compliance. Existing methods often ignore the temporal patterns in financial data, which are essential for understanding dynamic financial behaviors and detecting fraud. Moreover, they also treat companies as independent entities, overlooking the valuable interrelationships. To address these issues, we propose ACRF-RNN, a Recurrent Neural Network (RNN) with Attention-based Conditional Random Field (CRF) for fraud detection. Specifically, we use an RNN with a sliding window to capture temporal dependencies from historical data, and an attention-based CRF feature transformer to model inter-company relationships. This transforms raw financial data into optimized features, fed into a multi-layer perceptron for classification. Besides, we also use the focal loss to alleviate the class imbalance problem caused by rare fraudulent cases. This work presents a novel real-world dataset to evaluate the performance of ACRF-RNN. Extensive experiments show that ACRF-RNN outperforms the state-of-the-art methods by 15.28% in KS and 4.04% in Recall. Data and code are available at: https://github.com/XNetLab/ACRF-RNN.git. Xiaoguang Wang 0016, Chenxu Wang 0001, Luyue Zhang, Xiaole Wang, Mengqin Wang, Huanlong Liu, Tao Qin 0002 |
IJCAI | 7 |
| 2025 | Adversarial Propensity Weighting for Debiasing in Collaborative FilteringabstractDebiased recommendation focuses on alleviating the negative impact of various biases on recommendation quality to achieve fairer personalized recommendations. Current research mainly relies on propensity score estimation or causal inference methods to alleviate selection bias; at the same time, research on prevalence bias has proposed a variety of methods based on causal graphs and contrastive learning. However, these methods have shortcomings in dealing with unstable propensity score estimates, bias interactions, and decoupling of interest and bias signals, which limits the performance improvement of recommender systems. To this end, this paper proposes APWCF, a collaborative filtering debiased method that combines dynamic propensity modeling and adversarial learning. APWCF solves the problem of high variance in propensity scores through the dynamic propensity factor, and decouples user interests and bias signals through the adversarial learning to effectively remove multiple biases. Experiments show that APWCF significantly outperforms existing methods across various benchmark datasets from different domains. Compared with the current optimal baseline PDA, Recall@10 and NDCG@10 improve by 0.10%-5.42% and 1.01%-8.60% respectively. Kuiyu Zhu, Tao Qin 0002, Pinghui Wang |
IJCAI | 2 |
| 2025 | LRTHT: An Efficient Log Clustering Framework Based on Radix Tree and Hash Table
Yizhen Li, Tao Qin 0002, Jinzi Zou, Chenxu Wang 0001, Yuan-cheng Lu |
PAKDD (1) | 2 |
| 2025 | DAPEW: Towards robust collaborative filtering with graph contrastive learningabstractGraph Contrastive Learning (GCL) has shown excellent performance in Collaborative Filtering (CF), one of the most widely used techniques in efficient recommender systems . However, existing GCL-based CF methods suffer from node degree disparity, feature oversmoothing, difficulty in distinguishing hard negative samples, and semantic loss. To address these problems, this paper proposes a novel graph contrastive learning method for robust CF, named D egree- A ware P ropagation and E ntropy- W eighted contrastive loss (DAPEW). DAPEW introduces a degree-aware propagation mechanism to dynamically adjust the influence of initial embeddings, adjacency matrix products, and degree matrix products on the final embeddings, which can effectively handle node degree disparity and alleviate feature oversmoothing. DAPEW also designs an entropy-weighted contrastive loss, which introduces entropy weights to better distinguish hard negative samples and enhance the model’s discriminative ability and robustness. Experimental results show that DAPEW outperforms the existing GCL-based CF methods on several real-world datasets. Compared with existing GCL-based methods, DAPEW improves Recall@40 and NDCG@40 by 0.24% ∼ 25.88% and 0.14% ∼ 26.18% across four different datasets, respectively. Kuiyu Zhu, Tao Qin 0002, Haoxing Liu, Chenxu Wang 0001, Pinghui Wang |
Knowl. Based Syst. | 2 |
| 2025 | A Test Case Generation Scheduling Model Based on Running Event Interval CharacteristicsabstractSystem testing mainly focuses on finding bugs before online deployment. However, the system’s complexity increases due to a surge in the number of users, which makes scheduling the test case generation time more difficult, impeding testing fidelity and automatic testing. In this paper, we propose a data-driven framework for scheduling the test cases and improving testing fidelity. The developed framework involves two stages: 1) mining the running event interval characteristics from logs obtained from the actual environment in the first stage, and 2) utilizing those characteristics to develop a model for the test case generation time schedule in the second stage. First, we select the running event interval, which is the interval between two adjacent running events processed by the system, as the feature for analysis. Then, we analyze the interval’s distribution characteristics, observing that it adheres to the power-law distribution from multiple aspects. Second, based on the distribution characteristics obtained, we develop a composite event chain with associated relationships (CECAR) model to schedule test case generation time. Finally, we conduct a series of experiments using logs obtained from eight regions. The results show that, compared to baseline methods, the CECAR model reduces average errors in the generated quantity, power exponent, and specific granularity intervals by approximately 1% to 9% across multiple layers, demonstrating the effectiveness of the proposed method. We also made our dataset and code publicly available athttps://github.com/yizhenli-xjtu/CECAR. Note to Practitioners—System testing is a promising way to eliminate bugs and improve system stability before online deployment. If the testing method does not align with the system’s running characteristics in the actual environment, the testing fidelity and efficiency may be greatly reduced. Scheduling the test case generation time is one of the fundamental ways to address this challenge. The CECAR model developed in this paper can reproduce temporal characteristics and incorporate the concurrent relationships between different events, in turn, achieving the goal of fidelity improvement. Furthermore, based on the CECAR model, we can simulate arbitrary scenarios by combining different types of devices, which can greatly improve the testing’s flexibility and practicality. Yizhen Li, Tao Qin 0002, Jinzi Zou, Chao-Bo Yan, Xiaohong Guan |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2025 | TIRAGNN: Temporal and Implicit Relation-Aware Graph Neural Networks for Social RecommendationabstractSocial recommendation systems predict user preferences by using social relationships to address data sparsity and cold-start problems. Since social relations and user–item interactions can naturally be modeled as graph structures, graph neural networks (GNNs) have achieved significant success in social recommendation. However, most existing models fail to incorporate temporal information when modeling user–item interactions and rely solely on explicit social relationships to capture user influence, resulting in suboptimal performance. To address these problems, this article presents a temporal and implicit relation-aware graph neural network for social recommendation (TIRAGNN). Specifically, we model user and item representations using rating and temporal information from user–item interactions, integrating their relational influence within both the social graph and the constructed auxiliary graphs. Additionally, attention mechanisms are employed to model interaction sequences and aggregate relational influences, thereby enhancing the learning of user and item representations. Experimental results on two real-world datasets verify the superiority of TIRAGNN over state-of-the-art approaches. Chenxu Wang 0001, Dengdi Sun, Tao Qin 0002 |
IEEE Trans. Comput. Soc. Syst. | 5 |
| 2025 | Improving Representation Alignment and Uniformity With Siamese Graph Contrastive Learning for RecommendationabstractGraph contrastive learning (GCL) has shown excellent performance in recommender systems. However, GCL-based methods have two major problems. First, many methods use graph perturbations to create the contrastive view, this operation may compromise the structure of the original graph. Second, some algorithms pursue the uniformity of features, but neglect the alignment, which may lead to the features of similar samples in the feature space distancing from each other, thereby reducing the model’s ability to distinguish similar samples. To address these problems, we propose the SiamGCL framework, which consists of two modules: multilevel alignment (MLA) and enhanced uniformity (EU). MLA uses the siamese network structure to process user and item data in parallel and achieves model-level alignment through shared structure and parameters. In addition, we introduce alignment loss in the optimization process to further refine the alignment of user and item embeddings. EU employs noise modulated augmentation and batch normalization regulator as alternatives to traditional graph perturbation methods. This embedding perturbation strategy effectively prevents over-aggregation in the feature space, which can also enhance the uniformity of the embedding distribution. Experiments on five large-scale benchmark datasets show that SiamGCL improves Recall@20 and NDCG@20 by 0.97%–5.51% and 1.31%–2.73%, respectively, compared with the current state-of-the-art method. Kuiyu Zhu, Tao Qin 0002, Pinghui Wang, Xiaohong Guan |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2025 | Alignment-Guided Self-Supervised Learning for Diagram Question AnsweringabstractDiagram question answering (DQA), which is defined as answering natural language questions according to the visual diagram context, has attracted attention and has recently become a new benchmark for evaluating the complex reasoning ability of models. However, this reasoning task is extremely challenging because of the inclusion of abstract visual objects and specialized textual terms, as well as the complex relationships between them. The rarity of data caused by the high cost of annotation also makes large-scale deep models invalid for the DQA task. To address the above challenges, this paper proposes the cross-modal alignment-guided self-supervised learning model for DQA (CAS-DQA). Unlike previous works, the CAS-DQA model focuses on learning internal visual-textual object relationships, innovatively proposes an attention mechanism module based on object alignment, and effectively integrates cross-modal knowledge units for diagram understanding. In addition, the CAS-DQA model constructs two self-supervised learning (SSL) tasks via intermediate results of visual-textual object alignment. These two tasks exploit the unnoticed objects inside the diagram to fully and completely understand the diagram. They also effectively increase the amount of diagram question-answering data to address the challenge of data scarcity. To the best of our knowledge, the CAS-DQA model is the first to extend SSL strategies to the diagram question-answering task. We evaluate the CAS-DQA model on three different datasets. The results of extensive experiments show that our model significantly outperforms baselines on different scenarios and that the internal object alignment module and self-supervised tasks produce excellent results. Lingling Zhang 0005, Tao Qin 0002, Xinyu Zhang 0021, Jun Liu 0002 |
IEEE Trans. Multim. | 4 |
| 2025 | Knowledge Graph-Enhanced Masked Auto-Encoders for Recommendation SystemsabstractContrastive learning has received significant attention for its ability to improve the representation quality over limited labeled data, achieving notable advancements in Knowledge graphs (KGs) enhanced recommendation. However, the effectiveness of most contrastive methods heavily relies on manually designed data augmentation strategies, which often limit their generality across different datasets and downstream tasks, as well as their robustness to noise perturbations. To address these challenges, we propose a novel KG-enhanced recommendation framework namedKnowledgeGraph-enhancedMaskedAuto-Encoders forCollaborative Filtering (KG-CMAE) based on the masking-reconstruction paradigm, which employs two types of masked auto-encoders: the Knowledge Masked Auto-Encoder (KMAE) and the Collaborative Masked Auto-Encoder (CMAE), to adaptively extract informative self-supervised signals from KGs and user-item interactions. Specifically, KMAE employs the multi-head cross-attention mechanism to reflect the importance of neighboring nodes, and selectively masks and reconstructs the important connections with high structural consistency, thereby highlighting task-relevant knowledge. CMAE focuses on masking and reconstructing the interactions with high semantic relevance in the user-item interaction graph, and incorporates an enhanced decoder to better model the direct collaborative signals, such as user-user and item-item correlations, to mitigate the over-smoothing problem. Extensive experiments are conducted on three benchmark datasets under various settings, including noisy data, cold-start user recommendation, and long-tail item recommendation. The experimental results demonstrate the effectiveness, generality and robustness of the proposed KG-CMAE model compared to various baseline methods. Zhaoli Liu, Tao Qin 0002, Qindong Sun |
IEEE Trans. Serv. Comput. | 2 |
| 2024 | CoG-DQA: Chain-of-Guiding Learning with Large Language Models for Diagram Question AnsweringabstractDiagram Question Answering (DQA) is a challenging task, requiring models to answer natural language questions based on visual diagram contexts. It serves as a crucial basis for academic tutoring, technical support, and more practical applications. DQA poses significant challenges, such as the demand for domain-specific knowledge and the scarcity of annotated data, which restrict the applicability of large-scale deep models. Previous approaches have explored external knowledge integration through pretraining, but these methods are costly and can be limited by domain disparities. While Large Language Models (LLMs) show promise in question-answering, there is still a gap in how to cooperate and interact with the diagram parsing process. In this paper, we introduce the Chain-of-Guiding Learning Model for Diagram Question Answering (CoG-DQA), a novel framework that effectively addresses DQA challenges. CoG-DQA leverages LLMs to guide diagram parsing tools (DPTs) through the guiding chains, enhancing the precision of diagram parsing while introducing rich background knowledge. Our experimental findings reveal that CoG-DQA surpasses all comparison models in various DQA scenarios, achieving an average accuracy enhancement exceeding 5% and peaking at 11% across four datasets. These results underscore CoG-DQA's capacity to advance the field of visual question answering and promote the integration of LLMs into specialized domains. Lingling Zhang 0005, Longji Zhu, Tao Qin 0002, Kim-Hui Yap, Xinyu Zhang 0021, Jun Liu 0002 |
CVPR | 4 |
| 2024 | Multi-view cognition with path search for one-shot part labeling
Lingling Zhang 0005, Tao Qin 0002, Jun Liu 0002, Yifei Li 0006, Qianying Wang 0002 |
Comput. Vis. Image Underst. | 3 |
| 2024 | Distribution-aware hybrid noise augmentation in graph contrastive learning for recommendation
Kuiyu Zhu, Tao Qin 0002, Zhaoli Liu, Chenxu Wang 0001 |
Expert Syst. Appl. | 2 |
| 2024 | SynSem-ASTE: An Enhanced Multi-Encoder Network for Aspect Sentiment Triplet Extraction With Syntax and SemanticsabstractAspect Sentiment Triplet Extraction (ASTE) is an essential task in fine-grained opinion mining and sentiment analysis that involves extracting triplets consisting of aspect terms, opinion terms, and their associated sentiment polarities from texts. While prevailing approaches primarily adopt pipeline frameworks or unified tagging schemes for this task, these methods tend to either overlook syntactic structural information and inherent semantic features, or lack explicit mechanisms for integration of syntax and semantics among the triplets' elements. To overcome these shortcomings, we propose an Enhanced Multi-Encoder Network for ASTE with Syntax and Semantics (SynSem-ASTE). Our model innovatively incorporates syntactic information and semantic features derived from syntactic structures and attention weights, which is achieved through the design of a syntax encoder and a semantics encoder. Furthermore, we adopt a grid tagging scheme and an effective inference strategy to extract triplets simultaneously. Extensive evaluations on four benchmark datasets reveal that SynSem-ASTE not only achieves superior performance in terms of the primary metric F1-score, but also exhibits enhanced robustness against variations in model architecture. Lulin Liu, Tao Qin 0002, Yuankun Zhou, Chenxu Wang 0001, Xiaohong Guan |
IEEE Trans. Affect. Comput. | 2 |
| 2024 | GTCAlign: Global Topology Consistency-Based Graph AlignmentabstractGraph alignment aims to find correspondent nodes between two graphs. Most existing algorithms assume that correspondent nodes in different graphs have similar local structures. However, this principle may not apply to some real-world application scenarios when two graphs have different densities. Some correspondent node pairs may have very different local structures in these cases. Nevertheless, correspondent nodes are expected to have similar importance, inspiring us to exploit global topology consistency for graph alignment. This paper presents GTCAlign, an unsupervised graph alignment framework based on global topology consistency. An indicating matrix is calculated to show node pairs with consistent global topology based on a comprehensive centrality metric. A graph convolutional network (GCN) encodes local structural and attributive information into low-dimensional node embeddings. Then, node similarities are computed based on the obtained node embeddings under the guidance of the indicating matrix. Moreover, a pair of nodes are more likely to be aligned if most of their neighbors are aligned, motivating us to develop an iterative algorithm to refine the alignment results recursively. We conduct extensive experiments on real-world and synthetic datasets to evaluate the effectiveness of GTCAlign. The experimental results show that GTCAlign outperforms state-of-the-art graph alignment approaches. Chenxu Wang 0001, Peijing Jiang, Xiangliang Zhang 0001, Pinghui Wang, Tao Qin 0002, Xiaohong Guan |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2024 | Self-Supervised EEG Representation Learning for Robust Emotion RecognitionabstractEmotion recognition based on electroencephalography (EEG) is becoming a growing concern of researchers due to its various applications and portable devices. Existing methods are mainly dedicated to EEG feature representation and have made impressive progress. However, the problem of scarce labels restricts their further promotion. In light of this, we propose a self-supervised framework with contrastive learning for robust EEG-based emotion recognition, which can effectively leverage both readily available unlabeled EEG signals and labeled ones to learn highly discriminative EEG features. Firstly, we construct a specific pretext task according to the sequential non-stationarity of emotional EEG signals for contrastive learning, which aims at extracting pseudo-label information from all EEG data. Meanwhile, we propose a novel negative segment selection algorithm to reduce the noise of unlabeled data during the contrastive learning process. Secondly, to mitigate the overfitting issue induced by a small number of labeled samples during learning, we originate a loss function with label smoothing regularization that can guide the model to learn generalizable features. Extensive experiments over three benchmark datasets demonstrate the effectiveness and superiority of our model on EEG-based emotion recognition task. Besides, the generalization and robustness of the model have also been proved through sufficient experiments. Huan Liu 0012, Yuzhe Zhang 0003, Xuxu Chen, Dalin Zhang 0001, Rui Li 0073, Tao Qin 0002 |
ACM Trans. Sens. Networks | 6 |
| 2023 | Graph Contrastive Learning with Hybrid Noise Augmentation for Recommendation
Kuiyu Zhu, Tao Qin 0002, Zhouguo Chen, Jianwei Ding |
ADMA (4) | 2 |
| 2023 | Social Image-Text Sentiment Classification With Cross-Modal Consistency and Knowledge DistillationabstractSocial media sentiment analysis, which aims to evaluate the attitudes of online users based on their posts, has attracted significant research attention due to its successful application in the field of social media monitoring. It is a beneficial way to utilize multimodal information uploaded by users in order to improve sentiment classification ability. However, existing multimodal fusion-based approaches continue to face difficulties due to the issues of between-modality semantic inconsistency and missing modality. To address these issues, we propose a cross-modal consistency modeling-based knowledge distillation framework for image–text sentiment classification of social media data. Specifically, we design a hybrid curriculum learning strategy to measure the semantic consistency of multimodal data, then gradually train all image–text pairs from easy to hard, which can effectively handle the massive amounts of noise caused by inconsistencies between image and text data on social media. Moreover, in order to alleviate the problem of missing images in unimodal posts, we propose a privileged feature distillation method, in which the teacher model additionally considers images as privileged features, to transfer the visual knowledge to the student model, thereby enhancing the accuracy for text sentiment classification. Extensive experiments conducted over three real-world social media datasets demonstrate the effectiveness and superiority of the proposed multimodal sentiment analysis model. Huan Liu 0012, Jianping Fan 0007, Caixia Yan, Tao Qin 0002 |
IEEE Trans. Affect. Comput. | 5 |
| 2023 | EEG-Based Emotion Recognition With Emotion Localization via Hierarchical Self-AttentionabstractEmotion recognition based on electroencephalography (EEG) has attracted significant attention due to its wide range of applications, especially in Human-Computer Interaction(HCI). Previous research treats different segments of EEG signals uniformly, ignoring the fact that emotions are unstable and discrete during an extended period. In this paper, we propose a novel two-step spatial-temporal emotion recognition framework. First, considering that the human emotion has not only ”short-term continuity” but also ”long-term similarity”, we propose a hierarchical self-attention network to jointly model local and global temporal information, so as to localize most related segments and reduce the influence of noise at the temporal level. Second, in order to extract discriminative features at the spatial level to enhance the emotion recognition performance, we further employ the squeeze-and-excitation module (SE module) along with the channel correlation loss (CC-Loss) to select the most task-related channels. We also define a new task calledemotion localization, which aims to localize fragments with stronger emotions. We evaluate the proposed method on the proposed emotion localization task and typical emotion recognition task with three publicly available datasets, i.e., SEED, DEAP, and MAHNOB-HCI. The experimental results demonstrate that the proposed approach outperforms state-of-the-art methods. Yuzhe Zhang 0003, Huan Liu 0012, Dalin Zhang 0001, Xuxu Chen, Tao Qin 0002 |
IEEE Trans. Affect. Comput. | 5 |
| 2023 | SocialSift: Target Query Discovery on Online Social Media With Deep Reinforcement Learning
Changyu Wang, Pinghui Wang, Tao Qin 0002, Chenxu Wang 0001, Suhansanu Kumar, Xiaohong Guan, Jun Liu 0002, Kevin Chen-Chuan Chang |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | TPIPD: A Robust Model for Online VPN Traffic ClassificationabstractVPN has posed many difficulties for network security management. In this paper, we develop a robust method to classify the VPN traffic. Firstly, we investigate the VPN transmission process and find the turning packet interval (Named as TPI) is a valuable feature for VPN traffic classification. Then we employ the probability distribution of TPI to improve the robustness of classification process, which is named as TPIPD. Secondly, we evaluate our method using the ISCXVPN2016 dataset and find our method has higher classification accuracy compared with other related methods. We also find the distribution of the first few TPIs can be used to represent that of the entire TPIs of specific flow, thus our method can be used for online traffic classification. As TPIPD is a kind of probability feature, it is more robust than other traditional features. Finally, the experiments verify our methods can be used for mice flow identification. Yongwei Meng, Tao Qin 0002, Haonian Wang, Zhouguo Chen |
TrustCom | 2 |
| 2022 | SIRQU: Dynamic Quarantine Defense Model for Online Rumor Propagation ControlabstractRumors can spread very rapidly through online social networks (OSNs), leading to huge negative impact on human society. Hence, there is an urgent need to develop models that can minimize the spread of rumors. In this article, we propose a novel framework to improve the cost and efficiency of rumor propagation control. First, to reduce the impact of rumor controlling mechanism on users’ normal activities, we introduce a soft dynamic quarantine strategy into rumor propagation control and develop a new propagation model named susceptible-infected-removed-quarantined ignorants-quarantined spreaders (SIRQU) to model and block the rumor propagation in the network. Second, to further improve the control efficiency, we propose an influential node selection algorithm based on discrete particle swarm optimization with an evolutionary search strategy, and the controlling mechanism is only applied on the most influential nodes. Finally, we conduct a series of simulations and experiments on several public datasets and the dataset collected from Sina Weibo to validate the proposed method, and the results show that the proposed method outperfoms the related baseline algorithms. Zhaoli Liu, Tao Qin 0002, Qindong Sun, Shancang Li, Houbing Song, Zhouguo Chen |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2022 | Computer Science Diagram Understanding with Topology ParsingabstractDiagram is a special form of visual expression for representing complex concepts, logic, and knowledge, which widely appears in educational scenes such as textbooks, blogs, and encyclopedias. Current research on diagrams preliminarily focuses on natural disciplines such as Biology and Geography, whose expressions are still similar to natural images. In this article, we construct the first novel geometric type of diagrams dataset in Computer Science field, which has more abstract expressions and complex logical relations. The dataset has exhaustive annotations of objects and relations for about 1,300 diagrams and 3,500 question-answer pairs. We introduce the tasks of diagram classification (DC) and diagram question answering (DQA) based on the new dataset, and propose the Diagram Paring Net (DPN) that focuses on analyzing the topological structure and text information of diagrams. We use DPN-based models to solve DC and DQA tasks, and compare the performances to well-known natural images classification models and visual question answering models. Our experiments show the effectiveness of the proposed DPN-based models on diagram understanding tasks, also indicate that our dataset is more complex compared to previous natural image understanding datasets. The presented dataset opens new challenges for research in diagram understanding, and the DPN method provides a novel perspective for studying such data. Our dataset can be available from https://github.com/WayneWong97/CSDia. Lingling Zhang 0005, Yi Yang 0073, Tao Qin 0002, Jun Liu 0002 |
ACM Trans. Knowl. Discov. Data | 6 |
| 2022 | Heterogeneous Network Crawling: Reaching Target Nodes by Motif-Guided NavigationabstractWith numerous nodes on online heterogeneous networks, how to reach and extract target nodes of our specific interests is a pressing problem. In this paper, we propose a novel heterogeneous network crawler,MCrawl. It addresses the problem via iterative online heterogeneous network crawling by navigating its available APIs, starting from a set of target nodes, i.e., seed nodes. We are facing two challenges towards addressing the problem. First, to navigate within a vast network, how do we start from a small set of target nodes? In other words, which nodes in the “current frontier” and which direction shall we expand, to reach promising target nodes quickly? We propose motif-based crawling to exploit the complex structures and rich semantics of heterogeneous networks. Second, in many scenarios, we do not have a classifier to assess the quality of the harvested nodes and thus the motifs to expand. We develop a probabilistic inference framework to estimate the yield and harvest rates of motifs, achieving principled bootstrapping for crawling. Our experiment on real networks of MCrawl achieves significant margins over baselines. Changyu Wang, Kevin Chen-Chuan Chang, Pinghui Wang, Tao Qin 0002, Xiaohong Guan |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2020 | Credible seed identification for large-scale structural network alignment
Chenxu Wang 0001, Dong Qin, Xiapu Luo, Tao Qin 0002 |
Data Min. Knowl. Discov. | 6 |
| 2020 | Addressing the train-test gap on traffic classification combined subflow model with ensemble learning
Changyu Wang, Xiaohong Guan, Tao Qin 0002, Pinghui Wang, Yongwei Meng |
Knowl. Based Syst. | 3 |
| 2020 | A Two-Stagse Approach for Social Identity Linkage Based on an Enhanced Weighted Graph Model
Tao Qin 0002, Zhaoli Liu, Shancang Li, Xiaohong Guan |
Mob. Networks Appl. | 1 |
| 2020 | Symmetry Degree Measurement and its Applications to Anomaly DetectionabstractAnomaly detection is an important technique used to identify patterns of unusual network behavior and keep the network under control. Today, network attacks are increasing in terms of both their number and sophistication. To avoid causing significant traffic patterns and being detected by existing techniques, many new attacks tend to involve gradual adjustment of behaviors, which always generate incomplete sessions due to their running mechanisms. Accordingly, in this work, we employ the behavior symmetry degree to profile the anomalies and further identify unusual behaviors. We first proposed a symmetry degree to identify the incomplete sessions generated by unusual behaviors; we then employ a sketch to calculate the symmetry degree of internal hosts to improve the identification efficiency for online applications. To reduce the memory cost and probability of collision, we divide the IP addresses into four segments that can be used as keys of the hash functions in the sketch. Moreover, to further improve detection accuracy, a threshold selection method is proposed for dynamic traffic pattern analysis. The hash functions in the sketch are then designed using Chinese remainder theory, which can analytically trace the IP addresses associated with the anomalies. We tested the proposed techniques based on traffic data collected from the northwest center of CERNET (China Education and Research Network); the results show that the proposed methods can effectively detect anomalies in large-scale networks. Tao Qin 0002, Zhaoli Liu, Pinghui Wang, Shancang Li, Xiaohong Guan, Lixin Gao 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2019 | Blockchain-Based Digital Forensics Investigation Framework in the Internet of Things and Social SystemsabstractThe decentralized nature of blockchain technologies can well match the needs of integrity and provenances of evidences collecting in digital forensics (DF) across jurisdictional borders. In this paper, a novel blockchain-based DF investigation framework in the Internet of Things (IoT) and social systems environment is proposed, which can provide proof of existence and privacy preservation for evidence items examination. To implement such features, we present a block-enabled forensics framework for IoT, namely, IoT forensic chain (IoTFC), which can offer forensic investigation with good authenticity, immutability, traceability, resilience, and distributed trust between evidential entitles as well as examiners. The IoTFC can deliver a guarantee of traceability and track provenance of evidence items. Details of evidence identification, preservation, analysis, and presentation will be recorded in chains of block. The IoTFC can increase trust of both evidence items and examiners by providing transparency of the audit train. The use case demonstrated the effectiveness of the proposed method. Shancang Li, Tao Qin 0002, Geyong Min |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2018 | DeepMatching: A Structural Seed Identification Framework for Social Network AlignmentabstractNetwork alignment aims at finding a bijective mapping between nodes of two networks. Due to its wide application in various fields (e.g., Computer Vision, Data Management, Bioinformatics, and Privacy Protection), researchers have proposed many network alignment algorithms, most of which rely on a set of pre-mapped seeds. However, it is challenging to identify an initial credible set of seeds solely with structural information. In this paper, by exploiting the observation that a true mapping leads to a large portion of consistent edges among the mapped nodes, we formally define the credibility of a mapping as its deviation from a random one. This enables us to measure the credibility of an initial set of seeds. We also present DeepMatching which is a seed identification framework for social network alignment. First, we represent the nodes of the two mapping networks with their structural feature vectors by employing graph embedding techniques. Second, we obtain an initial mapping of the nodes based on the obtained vectors by leveraging point set registration methods. Third, we develop a heuristic algorithm to extract a credible set of seed from the initial mapping. Finally, we utilize the extracted seed set as input of an efficient propagation-based algorithm for large scale network alignment. We conduct extensive experiments to evaluate the performance of DeepMatching, and the results clearly demonstrate its effectiveness and the efficiency. Chenxu Wang 0001, Dong Qin, Xiapu Luo, Tao Qin 0002 |
ICDCS | 6 |
| 2018 | High Threat Alarms Mining for Effective Security Management: Modeling, Experiment and ApplicationabstractIntrusion Prevention System (IPS) is important for network security management as it can help the administrator by generating alarms corresponding to different attacks. But there are many false alarms due to their running mechanism, which greatly reduces its usability. In this paper, we develop a hierarchical framework to mine high threat alarms from raw massive logs. We first divide the raw alarms into two parts based on their attributes, the first part mainly include alarms from several kinds of serious attacks while others constitute the second part. To mine high threat alarms from the first part, we proposed a similar alarm mining method based on Choquet Integral to cluster and rank the results of clustering. The potential threats are mixed with many false alarms in the second part, to reduce effect from false alarms, we employ the frequent pattern mining algorithm to mine correlation rules and employ them to filter the false alarms. Following we qualify the threat degree of those alarms based on the features extracted from characteristics of alarms themselves. Experimental results based on the data collected from the campus network of Xi'an Jiaotong University verify the efficiency and accuracy of the developed methods. Based on the mining and ranking results, administrators can deal with the high threats with their limited time and energy to keep the network under control. Yongwei Meng, Tao Qin 0002 |
ISCC | 2 |
| 2018 | An efficient multi-feature SVM solver for complex event detection
Huan Liu 0012, Zhihui Li 0001, Tao Qin 0002, Lei Zhu 0002 |
Multim. Tools Appl. | 4 |
| 2017 | A traffic classification approach based on characteristics of subflows and ensemble learningabstractRecently, network traffic classification has attracted a great deal of attention among researchers. In this paper, we proposed a traffic classification approach based on characteristics of subflows and ensemble learning. Aiming at neutralization of unstable network environment as well as taking advantage of ensemble learning, we divided the traffic flows into different subflows in order to reduce the affection of time. Moreover, we develop truncation method on flows for real-time processing and an aggregation machine learning method based on accuracy of each classifier to different applications. Finally, the experimental results based on actual traffic traces collected from the campus network of Xian Jiaotong University verify the effectiveness of our methods. Changyu Wang, Xiaohong Guan, Tao Qin 0002 |
IM | 3 |
| 2017 | From online to offline: Charactering user's online music listening behavior for efficacious offline radio program arrangementabstractEfficacious radio music program arrangement with appropriate music played in the right time can greatly improve users' listening experience and attract more listeners for traditional radio stations. In this paper, we propose a method for efficacious offline radio program arrangement by mining users' online listening behavior characteristics in popular music websites. Firstly, we collect user's listening behavior trace of specific songs from two popular music sites in China. Secondly, the characteristics of basic listening behavior are analyzed and the concept of soaring degree is employed to quantify the popularity of a specific song. Different forecasting methods are also proposed for popularity prediction. Finally, based on the measurement and analysis results, we propose a simple and feasible strategy for offline program arrangement and the feedbacks from Shaanxi music station (FM 98.8) verify the efficiency and correctness of the proposed methods. Tao Qin 0002, Chenxu Wang 0001, Tao Yang 0006 |
ISCC | 1 |
| 2017 | Internet traffic classification based on expanding vector of flow
Jun Liu 0002, Tao Qin 0002, Haifei Li 0002 |
Comput. Networks | 3 |
| 2016 | A new model for nickname detection based on network structure and similarity propagationabstractUsers can participate in variety topics and express their opinions using different kinds of online applications, and the IDs (nickname) they used are usually virtual and difficult for finding the physical person. Which pose great challenges for network security management and user's online behavior supervision. Focus on this problem, we proposed methods for nickname detection based on the user's online connection structure and similarity propagation model. Firstly, we collected user's profile information from two popular online applications, sina microblog and RenRen network. Then we mark several matched pairs which belong to the same person from the applications with hardly manually effort, and those IDs are selected as seed set for nickname detection. Secondly, we proposed an nickname detection model based on the connection structure and similarity propagation. We selected one matched pairs from the seed set and obtain all their neighbors. Then we calculated the similarity of each pairs from the neighbor set and calculate their neighbors' connection similarity and neighbors' location similarity. If the similarity is bigger than a selected threshold, we claim they are matched pairs and insert them into the seed sets. On one hand, the correlation results can propagated based on the updated seed set. On the other hand, the computational complexity are greatly reduced as we only employ the neighbors' profiles to calculate the similarity. Experimental results verify the efficiency of the proposed method, which can lay a solid foundation for the online network management and user's behavior supervision. Zhaoli Liu, Tao Qin 0002, Xiaohong Guan, Tao Yang 0006 |
ISCC | 2 |
| 2016 | Modeling heterogeneous and correlated human dynamics of online activities with double Pareto distributions
Chenxu Wang 0001, Xiaohong Guan, Tao Qin 0002, Tao Yang 0006 |
Inf. Sci. | 3 |
| 2016 | CUFTI: Methods for core users finding and traffic identification in P2P systems
Tao Qin 0002 |
Peer-to-Peer Netw. Appl. | 1 |
| 2015 | Robust application identification methods for P2P and VoIP traffic classification in backbone networks
Tao Qin 0002, Zhaoli Liu, Xiaohong Guan |
Knowl. Based Syst. | 1 |
| 2015 | Analysis of user's behavior and resource characteristics for private trackers
Tao Qin 0002, Xiaohong Guan |
Peer-to-Peer Netw. Appl. | 2 |
| 2014 | A new connection degree calculation and measurement method for large scale network monitoring
Tao Qin 0002, Xiaohong Guan, Wei Li 0029, Pinghui Wang |
J. Netw. Comput. Appl. | 1 |
| 2014 | A New Sketch Method for Measuring Host Connection Degree DistributionabstractThe host connection degree distribution (HCDD) is an important metric for network security monitoring. However, it is difficult to accurately obtain the HCDD in real time for high-speed links with a massive amount of traffic data. In this paper, we propose a new sketch method to build a probabilistic traffic summary of a host's flows using a uniform Flajolet-Martin sketch combined with a small bitmap. To study its performance in comparison with previous sampling and sketch methods, we present a general model that encompasses all these methods. With this model, we compute the Cramér-Rao lower bounds and the variances of HCDD estimations. The theoretic analysis and numerical experimental results show that our sketch method is six times more accurate than state-of-the-art methods with the same memory usage. Pinghui Wang, Xiaohong Guan, Junzhou Zhao, Tao Qin 0002 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2013 | An in-depth measurement and analysis of popular private tracker systems in ChinaabstractIn recent years, a novel BitTorrent (BT) technology, Private Tracker (PT), has received extensive attentions in academia. Due to its huge amount of traffic volumes and population, it is essential to understand the characteristics for system management. In this paper, by using the data crawled from a number of well-known PT sites in China, we present an in-depth measurement and analysis on the PT system, in terms of resources characteristics and user behavior characteristics. Our analysis on the activity and age distribution of torrents reveals that the sequence of torrent activities follows a power-law distribution and the sequence of torrent ages follows a linear distribution. By investigating the characteristics of the core users and the resources in different PT sites, two observations have been made. Firstly, every PT site has a number of users with stable levels who contribute to the survivability and popularity of the site. Secondly, a large amount of redundancy exists among the distributed resources and many popular resources are duplicated among these sites, which leads to the quick spread of pirate resources. Based on the measurement results, we analyze the drawbacks of PT mechanisms and propose useful suggestions to promote the healthy development and performance of PT. Tao Qin 0002, Xiaohong Guan, Qiuzhen Huang |
ICC | 2 |
| 2013 | Inter-Swarm Content Distribution Among Private BitTorrent NetworksabstractPrivate BitTorrent (PT) is a new trend in Peer-to-Peer file sharing system, which provides high incentives for its users to seed after download by maintaining an upload-to-download ratio in the tracker for each registered community member. From the data we collected from six active PT sites, we discover that the population of both users and contents in any single PT site is much less than the public BitTorrent, and the intersection of content sets in different PTs is quite small. Based on this observation, we propose a content sharing/distribution framework among PTs (named CrossPT), as well as its sharing mechanism. In addition, we investigate the sharing strategy of the PT participants in CrossPT using game theory and the fetch strategy by modeling the scenario to a Neighbor Selection Problem (NSP). We prove NSP to be NP-complete and propose a heuristic algorithm to solve it. The evaluations with the input of crawled data from six PT sites demonstrate the efficiency of our mechanism. The content sizes of the six PT sites can be increased by 113.95%-438.46% with CrossPT. Also, the content distribution process can be done in less than one second, excluding the delivery time of the content itself. Chengchen Hu, Danfeng Shan, Yu Cheng 0003, Tao Qin 0002 |
IEEE J. Sel. Areas Commun. | 4 |
| 2012 | Behavior spectrum: An effective method for user's web access behavior monitoring and measurementabstractWith the number of Internet users and applications continues to grow, it becomes increasingly important to understand the users' behavior character for efficient network management and security monitoring. Web is one of the most popular applications, which can help users to obtain anything they want. In this paper, we propose a new framework to measure and monitor users' web access behavior character. Firstly, web services are divided into 12 types according to the content they provide. As user like to access different web services at different time instants, we proposed the behavior spectrum to describe the users' access character in an easily understandable way. Secondly, we employ three traffic features (number of connections, number of packets and number of bytes) to measure the intensity of users' access behavior. Based on the spectrums and features obtained, two kinds of user's access characters are analyzed: the characters at a particular time instant and the dynamic changing characters at continuous time points. Finally, we employ the Renyi entropy to cluster the users into different groups based on the spectrum and find the results are useful for network management. To verify the method, several actual traffic traces are collected from the Northwest Regional Center of CERNET (China Education and Research Network) and the experimental results prove the efficiency of the proposed methods. Tao Qin 0002, Wei Li 0029, Xiaohong Guan, Zhaoli Liu |
GLOBECOM | 1 |
| 2012 | Who are active? An in-depth measurement on user activity characteristics in sina microbloggingabstractThe emergence of online social network has changed the Internet users' online behaviors dramatically. Microblog, as an emerging online social network service, has attracted hundreds of millions of users to produce and share all kinds of content since its foundation. How to improve the Quality of Experience of the service and keep the users' passions for participant has become one of the most important problems confronted by the quick development of microblogging services. On the other hand, the QoE of a service influences and is influenced by its users' activities. In this paper, we analyzed the users' activities in Sina Weibo, a popular Chinese microblog website which has a profound influence on Chinese netizens. With the aim to find possible ways to improve users' Quality of Experience, we conduct a research to gain an in-depth understanding of what kinds of users are more active and what kinds of content are favored most. When we compare the distributions of the degrees of followers, we find Sina Weibo is still younger than Twitter, a famous worldwide microblogging service. Based on the following analysis, we find that users with around 300 friends in SinaWeibo are the most active users, which we believe is consistent with the Dunbar number. Furthermore, we find users in Sina Weibo are very keen on the content embedded with images and videos, which make up the majority of retweeted tweets. Through the analysis results, we confirm that inspiring users to follow others and improving the renderings of images and videos are the basic ways to improve the Quality of Experience. Chenxu Wang 0001, Xiaohong Guan, Tao Qin 0002, Wei Li 0029 |
GLOBECOM | 3 |
| 2011 | Monitoring abnormal network traffic based on blind source separation approach
Tao Qin 0002, Xiaohong Guan, Wei Li 0029, Pinghui Wang, Qiuzhen Huang |
J. Netw. Comput. Appl. | 1 |
| 2011 | A Data Streaming Method for Monitoring Host Connection Degrees of High-Speed LinksabstractDue to the massive amount of data in high-speed network traffic and the limit on processing capability, it is a great challenge to accurately measure and monitor network traffic over high-speed links online. A new data structure is presented in this paper for locating the hosts associated with large connection degrees or significant changes in connection degrees based on the reversible connection degree sketch to monitor anomalous network traffic. The reversible connection degree sketch builds a compact summary of host connection degrees efficiently and accurately. For each packet coming, it only needs to set several bits selected in a bit array by a group of hash functions. These hash functions are designed based on the Chinese Remainder Theorem so that the in-degree or out-degree associated with a given host can be accurately estimated. With this new data structure, we develop a new reverse sketch method for locating abnormal hosts. Although the reversible connection degree sketch does not preserve any host address information, we can analytically reconstruct the host addresses associated with large connection degrees or significant changes in connection degrees by a simple calculation purely based on the characteristics of the hash functions. Furthermore, a reinforced reversible connection degree sketch, the double connection degree sketch, is developed to reduce false positives which are commonly encountered in the sketch-based methods. A traffic monitoring system based on this double connection degree is developed to detect and classify the abnormal hosts associated with large connection degrees or significant changes in connection degrees. The experiments are conducted based on the actual network traffic and the testing results show that our method is accurate and efficient. Pinghui Wang, Xiaohong Guan, Tao Qin 0002, Qiuzhen Huang |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2010 | Dynamic Feature Analysis and Measurement for Large-Scale Network Traffic MonitoringabstractMeasuring and monitoring the changes of network traffic patterns in large-scale networks are crucial for effective network management. In this paper, we present a framework and method for detecting and measuring the dynamic changes of the pivotal traffic patterns. A bidirectional regional flow model is established to aggregate traffic packets and extract the traffic metrics and profiles. The characteristics of the regional flows are analyzed and interesting findings are obtained. A directed graph model is applied to describe the flow metrics and six flow features are extracted to capture the dynamic changes of the flow patterns. The measurements based on Renyi entropy are developed to quantitatively monitor these changes. The experimental results based on the actual network traffic data traces show that the method presented in this paper can capture the dynamic changes of pivotal traffic patterns effectively. Xiaohong Guan, Tao Qin 0002, Wei Li 0029, Pinghui Wang |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2009 | A New Data Streaming Method for Locating Hosts with Large Connection DegreeabstractLocating hosts with large connection degree is very important for monitoring anomalous network traffics. The in-degree (out-degree), defined as the number of distinct sources (destinations) that a network host is connected with (connects) during a given time interval. Due to massive amount of data in high speed network traffics and limit on processing capability, it is difficult to accurately locate hosts with large connection degree over high speed links on line. In this paper we present a new data streaming method for locating hosts with large connection degree based on the reversible connection degree sketch to monitor anomalous network traffics. The required memory space is small and constant, and more importantly the update/query complexity would not depend on the amount of data. The hash functions for data sketch are designed based on the remainder characteristics of the number theory so that in-degree/out-degree associated with a given host can be accurately estimated. Although the connection degree sketch does not preserve any host address information, we can analytically reconstruct the host addresses associated with large in-degree/out-degree by a simply equation purely based on the characteristics of the hash functions without using any host address information. This procedure is highly efficient since the computational time is constant and ignorable. Furthermore, this reversible connection degree sketch based method can be easily implemented in distributed systems. The experimental and testing results based on the actual network traffics show that the new method is truly accurate and efficient. Xiaohong Guan, Pinghui Wang, Tao Qin 0002 |
GLOBECOM | 3 |
| 2009 | Monitoring Abnormal Traffic Flows Based on Independent Component AnalysisabstractThe randomness of the network behaviors poses serious challenges for discovering the abnormal patterns in network traffic flows. This paper presents a method based on blind source separation approach for detecting abnormal traffic flows. It decomposes the network traffic into two components: the routine pattern and the abnormal pattern. The scale-space filter with adaptive scale is applied to filter the noise without affecting the main behavior patterns which can be used to form the abnormal traffic metrics and profiles. The zero-crossing method is applied to extract the stochastic behavior pulse widths and the largest width is selected as the scale space factor. In this way, the influence of the inherent randomness could be removed or greatly reduced. The extracted patterns of the routine behaviors imply the user's habit and the abnormal patterns are useful for discovering anomalous behaviors such as scanning, flooding and content distribution attacks. A salient feature of the method is that no supervised learning process is needed. This is a very important advantage since obtaining labeled samples in traffic monitoring is extremely difficult. Experimental results based on the datasets of an actual network show that this method is effective for monitoring anomaly traffic flows in the gigabytes traffic environment and the identification accuracy is above 95%. Tao Qin 0002, Xiaohong Guan, Wei Li 0029, Pinghui Wang |
ICC | 1 |