VLDB 2026 Research / reviewers in the wild / expert
Yanbing Liu 0007
dblp:84/4048-7
· DBLP profile ↗
70ranked-venue papers
4as first author
38since 2021 · last 2026
0000-0002-9653-073XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 35 · 19 since 2021Databases, data management, data science and information retrieval · 17 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 since 2021Human-computer interaction and ubiquitous computing · 7 · 7 since 2021Computer networks · 5 · 1 first-author · 2 since 2021Theory of computation · 2 · 2 first-authorSystems, architecture and hardware · 1 · 1 first-authorSecurity and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | OPERA: A Reinforcement Learning-Enhanced Orchestrated Planner-Executor Architecture for Reasoning-Oriented Multi-Hop RetrievalabstractRecent advances in large language models (LLMs) and dense retrievers have driven significant progress in retrieval-augmented generation (RAG). However, existing approaches face significant challenges in complex reasoning-oriented multi-hop retrieval tasks: 1) Ineffective reasoning-oriented planning: Prior methods struggle to generate robust multi-step plans for complex queries, as rule-based decomposers perform poorly on out-of-template questions. 2) Suboptimal reasoning-driven retrieval: Related methods employ limited query reformulation, leading to iterative retrieval loops that often fail to locate golden documents. 3) Insufficient reasoning-guided filtering: Prevailing methods lack the fine-grained reasoning to effectively filter salient information from noisy results, hindering utilization of retrieved knowledge. Fundamentally, these limitations all stem from the weak coupling between retrieval and reasoning in current RAG architectures. We introduce the Orchestrated Planner-Executor Reasoning Architecture (OPERA), a novel reasoning-driven retrieval framework. OPERA's Goal Planning Module (GPM) decomposes questions into sub-goals, which are executed by a Reason-Execute Module (REM) with specialized components for precise reasoning and effective retrieval. To train OPERA, we propose Multi-Agents Progressive Group Relative Policy Optimization (MAPGRPO), a novel variant of GRPO. Experiments on complex multi-hop benchmarks show OPERA's superior performance, validating both the MAPGRPO method and OPERA's design. Yanbing Liu 0007, Fangfang Yuan, Cong Cao 0001, Youbang Sun, Weizhuo Chen, Jianjun Li 0010, Zhiyuan Ma 0005 |
AAAI | 2 |
| 2026 | Two Streams, One Sarcasm: Orthogonal Expert Tuning for Holistic Multimodal Sarcasm UnderstandingabstractDiandian Guo, Cong Cao, Fangfang Yuan, Pin Xu, Cheng Hu, Zhicheng Zhang, Yu Liu, Yanbing Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Diandian Guo, Cong Cao 0001, Fangfang Yuan, Pin Xu, Yanbing Liu 0007 |
ACL (1) | 8 |
| 2026 | Trans-RAG: Query-Centric Vector Transformation for Secure Cross-Organizational Retrieval
Fangfang Yuan, Cong Cao 0001, Wenxuan Lu, Yanbing Liu 0007 |
DASFAA (4) | 7 |
| 2026 | FIRW: Frequency-injected Robust Watermarking for Latent Diffusion Models
Xiaokun Li, Fangfang Yuan, Cong Cao 0001, Majing Su, Yueshan Wang, Lei Jiang 0003, Yanbing Liu 0007 |
ICIC (2) | 7 |
| 2026 | MSD-NR: A Noise Rectification Method for Multimodal Sarcasm Detection
Xiangyu Tian, Diandian Guo, Cong Cao 0001, Yueshan Wang, Fangfang Yuan, Yanbing Liu 0007 |
ICIC | 6 |
| 2026 | Guided by LLM, Grounded in Targeted Vision: Dynamic Chain-of-Thought with Retrieval-Augmented In-Context Learning for Knowledge-Based Visual Question Answering
Yuling Yang, Weizhuo Chen, Cong Cao 0001, Fangfang Yuan, Yanbing Liu 0007 |
ICIC (19) | 6 |
| 2026 | TSS-RAG: Leveraging Teacher-Student Classroom Scenarios to Mitigate Hallucinations in Retrieval-Augmented Generation
Junqian Zhao, Tianhe Yu, Fangfang Yuan, Yanlong Zhou, Cong Cao 0001, Yanbing Liu 0007 |
ICIC | 7 |
| 2026 | MuVaC: A Variational Causal Framework for Multimodal Sarcasm Understanding in Dialogues
Diandian Guo, Fangfang Yuan, Cong Cao 0001, Xixun Lin, Chuan Zhou 0001, Hao Peng 0001, Yanan Cao 0006, Yanbing Liu 0007 |
WWW | 8 |
| 2025 | Towards More Reliable Chinese Spelling Correction: Fine-Grained Confidence Estimation Against Suboptimal Corrections
Chaodong Tong, Mingzhe Lu, Haimei Qin, Lei Jiang 0003, Yanbing Liu 0007 |
IEEE Big Data | 7 |
| 2025 | Dialogues Aspect-based Sentiment Quadruple Extraction via Structural Entropy Minimization PartitioningabstractDialogues Aspect-based Sentiment Quadruple Extraction (DiaASQ) aims to extract all target-aspect-opinion-sentiment quadruples from a given multi-round, multi-participant dialogue. Existing methods typically learn word relations across entire dialogues, assuming a uniform distribution of sentiment elements. However, we find that dialogues often contain multiple semantically independent sub-dialogues without clear dependencies between them. Therefore, learning word relationships across the entire dialogue inevitably introduces additional noise into the extraction process. To address this, our method focuses on partitioning dialogues into semantically independent sub-dialogues. Achieving completeness while minimizing these sub-dialogues presents a significant challenge. Simply partitioning based on reply relationships is ineffective. Instead, we propose utilizing a structural entropy minimization algorithm to partition the dialogues. This approach aims to preserve relevant utterances while distinguishing irrelevant ones as much as possible. Furthermore, we introduce a two-step framework for quadruple extraction: first extracting individual sentiment elements at the utterance level, then matching quadruples at the sub-dialogue level. Extensive experiments demonstrate that our approach achieves state-of-the-art performance in DiaASQ with much lower computational costs. Cong Cao 0001, Hao Peng 0001, Zhifeng Hao 0004, Lei Jiang 0003, Kongjing Gu, Yanbing Liu 0007, Philip S. Yu |
CIKM | 7 |
| 2025 | Multi-View Incongruity Learning for Multimodal Sarcasm DetectionabstractMultimodal sarcasm detection (MSD) is essential for various downstream tasks. Existing MSD methods tend to rely on spurious correlations. These methods often mistakenly prioritize non-essential features yet still make correct predictions, demonstrating poor generalizability beyond training environments. Regarding this phenomenon, this paper undertakes several initiatives. Firstly, we identify two primary causes that lead to the reliance of spurious correlations. Secondly, we address these challenges by proposing a novel method that integrate Multimodal Incongruities via Contrastive Learning (MICL) for multimodal sarcasm detection. Specifically, we first leverage incongruity to drive multi-view learning from three views: token-patch, entity-object, and sentiment. Then, we introduce extensive data augmentation to mitigate the biased learning of the textual modality. Additionally, we construct a test set, SPMSD, which consists potential spurious correlations to evaluate the the model’s generalizability. Experimental results demonstrate the superiority of MICL on benchmark datasets, along with the analyses showcasing MICL’s advancement in mitigating the effect of spurious correlation. Diandian Guo, Cong Cao 0001, Fangfang Yuan, Yanbing Liu 0007, Guangjie Zeng, Hao Peng 0001, Philip S. Yu |
COLING | 4 |
| 2025 | Emotion Transfer with Enhanced Prototype for Unseen Emotion Recognition in ConversationabstractCurrent Emotion Recognition in Conversation (ERC) research follows a closed-domain assumption.However, there is no clear consensus on emotion classification in psychology, which presents a challenge for models when it comes to recognizing previously unseen emotions in real-world applications.To bridge this gap, we introduce the Unseen Emotion Recognition in Conversation (UERC) task for the first time and propose ProEmoTrans, a solid prototype-based emotion transfer framework.This prototype-based approach shows promise but still faces key challenges: First, implicit expressions complicate emotion definition, which we address by proposing an LLM-enhanced description approach.Second, utterance encoding in long conversations is difficult, which we tackle with a proposed parameter-free mechanism for efficient encoding and overfitting prevention.Finally, the Markovian flow nature of emotions is hard to transfer, which we address with an improved Attention Viterbi Decoding (AVD) method to transfer seen emotion transitions to unseen emotions.Extensive experiments on three datasets show that our method serves as a strong baseline for preliminary exploration in this new area. Cong Cao 0001, Hao Peng 0001, Guanlin Wu, Zhifeng Hao 0004, Lei Jiang 0003, Yanbing Liu 0007, Philip S. Yu |
EMNLP | 7 |
| 2025 | ReTD: Reconstruction-Based Traceability Detection for Generated ImagesabstractThe objective of generated image traceability is to accurately identify and locate the source models. In this paper, we propose ReTD (Reconstruction-Based Traceability Detection), a generalized model for generated image traceability detection. Firstly, we use VAE to reconstruct images which are compared with the original ones to extract numerical distinguishing features. Secondly, we use Vision Transformer to learn the fine-grained distribution features to realize the generated image traceability classification. Finally, we conduct traceability experiments using images generated by ten GAN and Diffusion models. The experimental results demonstrate that ReTD only training a unified classifier improves accuracy by 9.4% compared to the state-of-the-art method. The ReTD-related code is availble at https://github.com/chenweizhuo/ReTD. Weizhuo Chen, Fangfang Yuan, Cong Cao 0001, Dakui Wang, Yanbing Liu 0007 |
ICASSP | 6 |
| 2025 | Counterfactual-Augmented Representation Learning based Event PredictionabstractAccurate prediction of future events holds significant importance for decision-makers. Current methods learn representations of past events from the observable graph structure to predict whether a future event will occur. However, these methods overlook the counterfactual scenarios, thus missing essential factors that could trigger future events. In this paper, we propose the Predicting Events with Counterfactual Augmentation Framework (PECF) to address this limitation. This is achieved by investigating whether deviations from observed events (i.e., counterfactual events) can affect the occurrence of the target future event. Specifically, first, we learn the representations of events through the temporal event graphs. Then, we instantiate causal models to represent the causal relationships between events. Finally, we generate counterfactual events and enhance event prediction accuracy through counterfactual-based augmentation. Experimental results demonstrate that our method outperforms current state-of-the-art methods on benchmark datasets. The code is available at https://github.com/hucheng-IIE/PECF. Fangfang Yuan, Cong Cao 0001, Guangjie Zeng, Yanbing Liu 0007, Hao Peng 0001, Philip S. Yu |
ICME | 6 |
| 2025 | CASD: Counterfactual Augmentation for Social Bot Detection on TwitterabstractSocial bot detection has become increasingly important with the rise of social media platforms, as social bots could be used for malicious activities such as spreading misinformation. Recent advances mainly utilize Graph Neural Networks (GNNs) for bot detection. However, these detection methods may overlook the issue of social bots’ disguise. An advanced bot can highly mimic human users and homogenize their features with human accounts, ultimately weakening the effectiveness of social bot detectors. In this paper, we propose CASD, a counterfactual-based social bot detection method. Specifically, we design a subgraph generation module to enhance node feature learning by reducing interference from irrelevant global information in sparse networks. Then, we apply counterfactual reasoning to simulate the disguising of bots and identify their distinguishable features. Finally, we combine the counterfactual feature with its raw counterpart for robust bot detection. Experimental results on multiple datasets demonstrate that our method outperforms existing state-of-the-art (SOTA) methods. Our code is publicly available at1. Pin Xu, Fangfang Yuan, Yueshan Wang, Diandian Guo, Cong Cao 0001, Yanbing Liu 0007 |
ICME | 6 |
| 2025 | T-T: Table Transformer for Tagging-based Aspect Sentiment Triplet ExtractionabstractAspect sentiment triplet extraction (ASTE) aims to extract triplets composed of aspect terms, opinion terms, and sentiment polarities from given sentences. The table tagging method is a popular approach to addressing this task, which encodes a sentence into a 2-dimensional table, allowing for the tagging of relations between any two words. Previous efforts have focused on designing various downstream relation learning modules to better capture interactions between tokens in the table, revealing that a stronger capability in relation capture can lead to greater improvements in the model. Motivated by this, we attempt to directly utilize transformer layers as downstream relation learning modules. Due to the powerful semantic modeling capability of transformers, it is foreseeable that this will lead to excellent improvement. However, owing to the quadratic relation between the length of the table and the length of the input sentence sequence, using transformers directly faces two challenges: overly long table sequences and unfair local attention interaction. To address these challenges, we propose a novel Table-Transformer (T-T) for the tagging-based ASTE method. Specifically, we introduce a stripe attention mechanism with a loop-shift strategy to tackle these challenges. The former modifies the global attention mechanism to only attend to a 2-dimensional local attention window, while the latter facilitates interaction between different attention windows. Extensive and comprehensive experiments demonstrate that the T-T, as a downstream relation learning module, achieves state-of-the-art performance with lower computational costs. Chaodong Tong, Cong Cao 0001, Hao Peng 0001, Qian Li 0033, Guanlin Wu, Lei Jiang 0003, Yanbing Liu 0007, Philip S. Yu |
IJCAI | 8 |
| 2025 | FairCDR: Transferring Fairness and User Preferences for Cross-Domain RecommendationabstractCross-domain recommendation (CDR) has gained significant attention for its ability to address data sparsity issue. However, most existing CDR methods focus primarily on improving recommendation accuracy while largely overlooking fairness considerations, which can lead to biased outcomes and unfair treatment of different user groups. To solve this critical problem, we investigate whether fairness can be transferred from the source domain to the target domain. Our analysis suggests that fairness can be effectively transferred if the fairness of the source domain is ensured and the distributions of the source and target domains are well aligned. Based on this, we propose the FairCDR, a novel framework that can achieve the knowledge transfer of fairness and user preferences simultaneously. FairCDR owns two phases: single-domain fairness guarantee and inter-domain distribution alignment. In the first phase, we employ an adversarial learning-based recommender (ALR) to disentangle user preferences from sensitive attributes in the source domain. In the second phase, we introduce a new mutual learning-based diffusion model (MLDiff), which engages in mutual learning with ALR to progressively align the distributions of the source and target domains. This improves ALR's adaptability to distribution shifts, ultimately ensuring fairness and recommendation performance in the target domain. Extensive experiments on multiple real-world cross-domain datasets demonstrate that FairCDR surpasses existing strong baselines in both fairness and recommendation quality. Yongxuan Wu, Yang Aron Liu, Xixun Lin, Yanan Cao 0001, Lixin Zou, Yanmin Shang, Yanbing Liu 0007 |
KDD (2) | 8 |
| 2025 | RepGuard: Adaptive Feature Decoupling for Robust Backdoor Defense in Large Language ModelsabstractBackdoor attacks pose a significant threat to large language models (LLMs) by embedding malicious triggers that manipulate model behavior. However, existing defenses primarily rely on prior knowledge of backdoor triggers or targets and offer only superficial mitigation strategies, thus struggling to fundamentally address the inherent reliance on unreliable features. To address these limitations, we propose a novel defense strategy, \textit{RepGuard}, that strengthens LLM resilience by adaptively separating abnormal features from useful semantic representations, rendering the defense agnostic to specific trigger patterns. Specifically, we first introduce a dual-perspective feature localization strategy that integrates local consistency and sample-wise deviation metrics to identify suspicious backdoor patterns. Based on this identification, an adaptive mask generation mechanism is applied to isolate backdoor-targeted shortcut features by decomposing hidden representations into independent spaces, while preserving task-relevant semantics. With a multi-objective optimization framework, our method can inherently mitigates backdoor attacks. Across \textit{Target Refusal} and \textit{Jailbreak} tasks under four types of attacks, RepGuard consistently reduced the attack success rate on poisoned data by nearly 80\% on average, while maintaining near-original task performance on clean data. Extensive experiments demonstrate that RepGuard provides a scalable and interpretable solution for safeguarding LLMs against sophisticated backdoor threats. Jie Zhang 0050, Yanbing Liu 0007, Yunpeng Li 0006, Jinta Weng, Yue Hu 0002 |
NeurIPS | 3 |
| 2025 | See Better, Say Better: Vision-Augmented Decoding for Mitigating Hallucinations in Large Vision-Language Models
Xinyi Sun, Diandian Guo, Cong Cao 0001, Fangfang Yuan, Dakui Wang, Yanbing Liu 0007 |
NLPCC (1) | 6 |
| 2024 | Multi-Level Graph Convolutional Network for Document Information ExtractionabstractDocument information extraction aims to identify and extract entities and other essential information from documents. Its performance is significantly dependent on the learned relationship between the text and its corresponding layout, which is not fully exploited by existing methods. Therefore, we propose a multi-level graph convolutional network to improve the capability for modeling document information. Firstly, we integrate text embeddings with corresponding image and layout features as text features, which are used to construct a text-level graph to extract semantic relationships as word features through graph convolutions. We then construct a layout-level graph using region layout features as nodes, extracting structure relationships through further graph convolutions. This multi-level graph structure allows our model to fuse fine-grained word features with coarse-grained region features for effective sequence labeling. Comprehensive experiments on various datasets consistently demonstrate the effectiveness of our method, and the results show that our model outperforms the state-of-the-art methods with the aid of multi-level features. Fangfang Yuan, Dakui Wang, Cong Cao 0001, Yanbing Liu 0007 |
ICTAI | 7 |
| 2024 | Efficient One-Shot Pruning of Large Language Models with Low-Rank ApproximationabstractModel pruning, as an effective method for compressing large language models (LLMs), has recently attracted considerable attention in the field of natural language processing. However, existing LLM pruning methods have two main drawbacks: (1) Iterative pruning for LLMs with over a billion parameters requires retraining, which leads to significant pruning costs. (2) LLMs Pruning is formalized as a weight reconstruction problem that necessitates second-order information, incurring expensive computations. To address these issues, we propose a novel pruning method named Eplra: efficient one-shot pruning of large language models with low-rank approximation, which efficiently identifies sparse networks in LLMs. Specifically, we design a novel pruning metric based on input activations for the rapid one-shot compression of LLMs. We first incorporate input activations into the calculation of weight importance to promote precise pruning of low-priority weights. Then, we perform local weight comparisons across each output of linear layers to induce uniform sparsity. Next, we expand Eplra into semi-structured pruning patterns to accommodate various acceleration scenarios. Finally, we employ low-rank parametrized update matrices to fine-tune the pruned model, facilitating a swift recovery of model performance. Experimental results on various language benchmark datasets demonstrate that Eplra outperforms the state-of-the-art methods. Yangyan Xu, Cong Cao 0001, Fangfang Yuan, Rongxin Mi, Nannan Sun, Dakui Wang, Yanbing Liu 0007 |
SMC | 7 |
| 2024 | Enhancing GPT-3.5 for Knowledge-Based VQA with In-Context Prompt Learning and Image CaptioningabstractTraditional visual question answering (VQA) often falls short as merely relying on image information is insufficient to answer given questions. Therefore, Knowledge-Based Visual Question Answering (KB-VQA) has emerged. Typically, KB-VQA involves first retrieving knowledge from external knowledge bases, then using the retrieved knowledge in conjunction with the understanding of visual content for joint reasoning to predict answers. However, current models often suffer from weak visual perception capabilities when processing image information. Additionally, due to the incompleteness of external knowledge bases, retrieved knowledge may contain noise or even irrelevant information. Moreover, the re-embedding of knowledge text features during the model's reasoning process may deviate from the original meanings in the knowledge base. To address these challenges, we propose a method for Knowledge-Based Visual Question Answering (KB-VQA) using GPT-3.5, leveraging image captions and in-context prompts. We utilize an advanced captioning model to convert images into accurate textual representations, enhancing the large language model's understanding of visual information. Moreover, we eliminate the need for additional knowledge bases by directly employing GPT-3.5 as a knowledge base for knowledge retrieval and generate logically consistent text during inference to predict answers. Furthermore, we enhance GPT-3.5's question-answering capability for VQA through in-context prompt learning. Experiments on the public OK-VQA dataset demonstrate the superior performance of our model. Yuling Yang, Cong Cao 0001, Fangfang Yuan, Dakui Wang, Yanbing Liu 0007 |
SMC | 6 |
| 2024 | RDGCN: Reinforced Dependency Graph Convolutional Network for Aspect-based Sentiment AnalysisabstractAspect-based sentiment analysis (ABSA) is dedicated to forecasting the sentiment polarity of aspect terms within sentences. Employing graph neural networks to capture structural patterns from syntactic dependency parsing has been confirmed as an effective approach for boosting ABSA. In most works, the topology of dependency trees or dependency-based attention coefficients is often loosely regarded as edges between aspects and opinions, which can result in insufficient and ambiguous syntactic utilization. To address these problems, we propose a new reinforced dependency graph convolutional network (RDGCN) that improves the importance calculation of dependencies in both distance and type views. Initially, we propose an importance calculation criterion for the minimum distances over dependency trees. Under the criterion, we design a distance-importance function that leverages reinforcement learning for weight distribution search and dissimilarity control. Since dependency types often do not have explicit syntax like tree distances, we use global attention and mask mechanisms to design type-importance functions. Finally, we merge these weights and implement feature aggregation and classification. Comprehensive experiments show the effectiveness of the criterion and importance functions. RDGCN yields excellent analysis results. Xusheng Zhao, Hao Peng 0001, Qiong Dai, Huailiang Peng, Yanbing Liu 0007, Qinglang Guo, Philip S. Yu |
WSDM | 6 |
| 2023 | NPGraph: An Efficient Graph Computing Model in NUMA-Based Persistent Memory Systems
Baoke Li, Cong Cao 0001, Fangfang Yuan, Yuling Yang, Majing Su, Yanbing Liu 0007, Jianhui Fu |
CollaborateCom (2) | 6 |
| 2023 | Few-shot Malicious Domain Detection on Heterogeneous Graph with Meta-learningabstractThe Domain Name System (DNS), one of the essential basic services on the Internet, is often abused by attackers to launch various cyber attacks, such as phishing and spamming. Researchers have proposed many machine learning-based and deep learning-based methods to detect malicious domains. However, these methods rely on a large-scale dataset with labeled samples for model training. The fact is that the labeled domain samples are limited in the real-world DNS dataset. In this paper, we propose a few-shot malicious domain detection model named MetaDom, which employs a meta-learning algorithm for model optimization. Specifically, We first model the DNS scenario as a heterogeneous graph to capture richer information by analysing the complex relations among domains, IP addresses and clients. Then, we learn the domain representations with a heterogeneous graph neural network on the DNS HG. Finally, considering that only few labeled data are available in the real-world DNS scenario, a meta-learning algorithm with knowledge distillation is introduced to optimize the model. Extensive experiments on the real DNS dataset show that MetaDom outperforms other state-of-the-art methods. Fangfang Yuan, Cong Cao 0001, Majing Su, Dakui Wang, Yanbing Liu 0007 |
CSCWD | 6 |
| 2023 | Curvature-Driven Knowledge Graph Embedding for Link PredictionabstractKnowledge Graph Embedding (KGE) aims to learn how to represent the low-dimensional vectors for entities and relations based on the observed triplets in knowledge graph. Most of the existing models use simple structural features, such as node degrees and directed edges, and pay little attention to advanced inherent information of structured knowledge. In this paper, we propose CD-GCN, a curvature-driven KGE method for link prediction. Specifically, we first apply Ricci curvature to knowledge graph. Then, we use curvature information to drive the state update, which aims to further exploit the graph-structured information. Finally, we use a ConvE scoring function to output the link prediction results. Through extensive experiments on public datasets FB15k-237 and WN18RR, CD-GCN has achieved state-of-the-art results compared with all baseline models. Diandian Guo, Majing Su, Cong Cao 0001, Fangfang Yuan, Yanbing Liu 0007, Jianhui Fu |
CSCWD | 6 |
| 2023 | MetaBERT: Collaborative Meta-Learning for Accelerating BERT InferenceabstractEarly exit methods are used to accelerate inference in pre-trained language models and maintain competitive performance on resource-constrained devices. However, existing methods for training early exit classifiers suffer from the problem of poor classifier representations in different layers, leading to difficulties in adapting to diverse natural language processing tasks. To address this issue, we propose MetaBERT: collaborative Meta-learning for accelerating BERT inference. The main goal of MetaBERT is to train early exit classifiers through collaborative meta-learning, in which case, few gradient updates can be quickly adapted to new tasks. Moreover, this novel meta-training approach produces good generalization performance, thus achieving an effective balance between the inference result and efficiency. Extensive experimental results show that our approach outperforms previous training methods by a large margin, and achieves state-of-the-art results compared to other competitive models. Yangyan Xu, Fangfang Yuan, Cong Cao 0001, Majing Su, Dakui Wang, Yanbing Liu 0007 |
CSCWD | 7 |
| 2023 | Malicious Domain Detection Based on Self-supervised HGNNs with Contrastive Learning
Zhiping Li, Fangfang Yuan, Cong Cao 0001, Majing Su, Yuhai Lu, Yanbing Liu 0007 |
ICANN (3) | 6 |
| 2023 | EDDVPL: A Web Attribute Extraction Method with Prompt Learning
Yuling Yang, Jiali Feng, Baoke Li, Fangfang Yuan, Cong Cao 0001, Yanbing Liu 0007 |
ICONIP (14) | 6 |
| 2023 | RegexClassifier: A GNN-Based Recognition Method for State-Explosive Regular ExpressionsabstractRegular expression (regex) matching technology has been widely used in various applications. For the sake of low time complexity and stable performance, Deterministic Finite Automaton (DFA) has become the first choice to perform fast regular expression matching. However, DFA has the state explosion problem, that is, the number of DFA states may increase exponentially while compiling some specific regexes to DFA. The huge memory consumption restricts its practical applications. A lot of works have addressed the DFA state explosion problem; however, none has met the requirements of fast recognition and small memory image. In this paper, we proposed RegexClassifier to recognize state-explosive regexes intelligently and efficiently. It firstly transforms regexes into Non-deterministic Finite Automatons(NFAs), then uses Graph Neural Network(GNN) models to classify NFAs in order to recognize regexes that may cause DFA state explosion. Experiments on typical rule sets show that the classification accuracy of the proposed model is up to 98%. Yuhai Lu, Fangfang Yuan, Cong Cao 0001, Yanbing Liu 0007 |
ISCC | 6 |
| 2023 | Robust Malicious Domain Detection Against Adversarial Attacks on Heterogeneous GraphabstractDomain Name System (DNS) is a crucial infrastructure of the Internet, yet it is also a primary medium for disseminating illicit information. Researchers have proposed numerous methods to detect malicious domains, among which heterogeneous graph (HG) based models have demonstrated good performance. However, their success may also motivate attackers to defeat HG based models in order to evade detection. In this paper, we propose a novel malicious domain detection model named RoDom, which is robust against adversarial attacks on HG. Firstly, we introduce different perturbations to construct multiple attacked graphs, which are designed to simulate different types of adversarial attacks on the HG. Secondly, we design a discriminator to perform robust representation learning on the HG by discriminating the original graph from attacked graphs. Finally, we introduce a classification selector to further improve the model's robustness by automatically combining domain representations of multiple HGs for domain classification. The experimental results show that RoDom out-performs other state-of-the-art methods and exhibits stronger robustness against adversarial attacks on the HG. Zhiping Li, Fangfang Yuan, Dakui Wang, Cong Cao 0001, Yanbing Liu 0007 |
SMC | 7 |
| 2022 | Heterogeneous Graph Attention Network for Malicious Domain Detection
Zhiping Li, Fangfang Yuan, Yanbing Liu 0007, Cong Cao 0001, Fang Fang 0009, Jianlong Tan |
ICANN (2) | 3 |
| 2022 | DOM2R-Graph: A Web Attribute Extraction Architecture with Relation-Aware Heterogeneous Graph Transformer
Jiali Feng, Cong Cao 0001, Fangfang Yuan, Zhiping Li, Yanbing Liu 0007, Jianlong Tan |
ICONIP (1) | 6 |
| 2022 | Malicious Domain Detection with Heterogeneous Graph Propagation Network
Fangfang Yuan, Yanbing Liu 0007, Cong Cao 0001, Jianlong Tan |
WASA (1) | 3 |
| 2021 | Deep Differential Amplifier for Extractive SummarizationabstractRuipeng Jia, Yanan Cao, Fang Fang, Yuchen Zhou, Zheng Fang, Yanbing Liu, Shi Wang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Ruipeng Jia, Yanan Cao 0001, Fang Fang 0009, Zheng Fang 0002, Yanbing Liu 0007, Shi Wang 0002 |
ACL/IJCNLP (1) | 6 |
| 2021 | Malicious Domain Detection on Imbalanced Data with Deep Reinforcement Learning
Fangfang Yuan, Teng Tian, Yanmin Shang, Yuhai Lu, Yanbing Liu 0007, Jianlong Tan |
ICONIP (4) | 5 |
| 2021 | Knowledge-Based Diverse Feature Transformation for Few-Shot Relation Classification
Yubao Tang, Zhezhou Li, Cong Cao 0001, Fang Fang 0009, Yanan Cao 0001, Yanbing Liu 0007, Jianhui Fu |
KSEM | 6 |
| 2021 | RLINK: Deep reinforcement learning for user identity linkageabstractAbstract User identity linkage is a task of recognizing the identities of the same user across different social networks (SN). Previous works tackle this problem via estimating the pairwise similarity between identities from different SN, predicting the label of identity pairs or selecting the most relevant identity pair based on the similarity scores. However, most of these methods fail to utilize the results of previously matched identities, which could contribute to the subsequent linkages in following matching steps. To address this problem, we transform user identity linkage into a sequence decision problem and propose a reinforcement learning model to optimize the linkage strategy from the global perspective. Our method makes full use of both the social network structure and the history matched identities, meanwhile explores the long-term influence of processing matching on subsequent decisions. We conduct extensive experiments on real-world datasets, the results show that our method outperforms the state-of-the-art methods. Yanan Cao 0001, Qian Li 0003, Yanmin Shang, Yangxi Li, Yanbing Liu 0007, Guandong Xu |
World Wide Web | 6 |
| 2020 | Type-Aware Anchor Link Prediction across Heterogeneous Networks Based on Graph Attention NetworkabstractAnchor Link Prediction (ALP) across heterogeneous networks plays a pivotal role in inter-network applications. The difficulty of anchor link prediction in heterogeneous networks lies in how to consider the factors affecting nodes alignment comprehensively. In recent years, predicting anchor links based on network embedding has become the main trend. For heterogeneous networks, previous anchor link prediction methods first integrate various types of nodes associated with a user node to obtain a fusion embedding vector from global perspective, and then predict anchor links based on the similarity between fusion vectors corresponding with different user nodes. However, the fusion vector ignores effects of the local type information on user nodes alignment. To address the challenge, we propose a novel type-aware anchor link prediction across heterogeneous networks (TALP), which models the effect of type information and fusion information on user nodes alignment from local and global perspective simultaneously. TALP can solve the network embedding and type-aware alignment under a unified optimization framework based on a two-layer graph attention architecture. Through extensive experiments on real heterogeneous network datasets, we demonstrate that TALP significantly outperforms the state-of-the-art methods. Yanmin Shang, Yanan Cao 0001, Yangxi Li, Jianlong Tan, Yanbing Liu 0007 |
AAAI | 6 |
| 2020 | DistilSum: : Distilling the Knowledge for Extractive SummarizationabstractA popular choice for extractive summarization is to conceptualize it as sentence-level classification, supervised by binary labels. While the common metric ROUGE prefers to measure the text similarity, instead of the performance of classifier. For example, BERTSUMEXT, the best extractive classifier so far, only achieves a precision of 32.9% at the top 3 extracted sentences ([email protected]) on CNN/DM dataset. It is obvious that current approaches cannot model the complex relationship of sentences exactly with 0/1 targets. In this paper, we introduce DistilSum, which contains teacher mechanism and student model. Teacher mechanism produces high entropy soft targets at a high temperature. Our student model is trained with the same temperature to match these informative soft targets and tested with temperature of 1 to distill for ground-truth labels. Compared with large version of BERTSUMEXT, our experimental result on CNN/DM achieves a substantial improvement of 0.99 ROUGE-L score (text similarity) and 3.95 [email protected] score (performance of classifier). Our source code will be available on Github. Ruipeng Jia, Yanan Cao 0001, Haichao Shi, Fang Fang 0009, Yanbing Liu 0007, Jianlong Tan |
CIKM | 5 |
| 2020 | Data Augmentation for Insider Threat Detection with GANabstractIn insider threat detection domain, the datasets are highly imbalanced, where the number of user's normal behavior is higher than that of insider's anomalous behavior. A direct approach to handle the class imbalance problem is using data augmentation on the minority class. Existing data augmentation methods mainly produce synthetic samples according with the linear operation based on samples of the minority class. Hence, these methods just focus on local information which leads to the unitarily of the synthetic samples, resulting in overfitting. To enrich the diversity of the synthetic samples, we propose a deep adversarial insider threat detection (DAITD) framework using the Generative Adversarial Networks (GAN) to approximate the true anomalous behavior distribution. Specifically, we first obtain anomalous user behavior representations from the anomalous behavior data (minority class), and then use the generator of the GAN to model the actual anomalous behavior distribution, use the discriminator of the GAN to distinguish whether the synthetic sample from the generator is real or not. In this way, our method is able to generate high quality synthetic samples that are close to the anomalous user behavior. Experimental results show that the DAITD framework outperforms other comparative inside threat detection algorithms. Fangfang Yuan, Yanmin Shang, Yanbing Liu 0007, Yanan Cao 0001, Jianlong Tan |
ICTAI | 3 |
| 2020 | Enhancing Textual Representation for Abstractive Summarization: Leveraging Masked DecoderabstractFor existing models of abstractive summarization, the paradigm of autoregressive decoder inherently prefers relying on former tokens and the prediction error will propagate subsequently. To effectively eliminate the errors, we need a way to remodeling dependency during text generation. In this paper, we introduce MDSumma (as shorthand for Masked Decoder for Summarization), which masks partial tokens in decoder, aiming to alleviate the over-reliance on the antecedent. Moreover, with further facilitating the flexibility and diversity of textual representation, we employ a variational autoencoder model, sampling continuous latent variables from the probability distribution to explicitly model underlying semantics of the target summaries. Our architecture gives good balance between encoder contextual representation and decoder prediction, sidestepping the gap between training and inference. Experimental results on three benchmark datasets validate the effectiveness that our proposed method significantly outperforms the existing state-of-the-art approaches both on ROUGE and diversity scores. Ruipeng Jia, Yannan Cao, Fang Fang 0009, Jinpeng Li 0003, Yanbing Liu 0007, Pengfei Yin |
IJCNN | 5 |
| 2020 | Enhancing Pre-trained Language Representation for Multi-Task Learning of Scientific SummarizationabstractThis paper aims to extract summarization and keywords from scientific articles simultaneously, while abstract extraction (AE) and key extraction (KE) are considered as auxiliary tasks to each other. For the data scarcity in scientific AE and KE tasks, we propose a multi-task learning framework which uses huge unlabeled data to learn scientific language representation (pre-training) and uses smaller annotated data to transfer the learned representation to AE and KE (fine-tuning). Although the pre-trained language model performs well in universal natural language tasks, its capacity still has a margin of improvement for specific tasks. Inspired by this intuition, we use another two tasks keyword masking and key sentence prediction before the fine-tuning phase to enhance the language representation for AE and KE. This language representation enhancing stage uses the same labeled data but different optimization objectives with the fine-tuning phase. In order to evaluate our model, we develop and release a high-quality annotated corpus for scientific papers with keywords and abstract. We conduct comparative experiments on this dataset, and experimental results show that our multi-task learning framework achieves the state-of-the-art performance, proving the effectiveness of the language model enhancing mechanism. Ruipeng Jia, Yannan Cao, Fang Fang 0009, Jinpeng Li 0003, Yanbing Liu 0007, Pengfei Yin |
IJCNN | 5 |
| 2020 | A Hybrid Paper Recommendation Method by Using Heterogeneous Graph and MetadataabstractThe amount of academic articles in digital libraries is increasing exponentially. This growth of scientific papers' growth made it difficult for researchers to obtain related papers from their queries. Recommendation systems can help them resolve the problem of information overload. However, existing paper recommender methods generally rely on the simple citation network, which ignores the semantic of papers and has the problem of cold start. In this paper, a hybrid paper recommendation approach AMHG is proposed which is based on a multi-level citation heterogeneous graph. Unlike existing works which only use the reference relationship, we consider the same or similar authors' papers to alleviate the cold start problem of zero-citation and newly published papers. Besides, the metadata information of papers is also incorporated into a representation model to generate better recommender results to alleviate the cold start problem. We use the authors' influence factors to reorder the candidate list outputting by MLP to obtain high-quality articles. Through experiments, we compare our model with several methods on the DBLP-REC dataset to demonstrate that AMHG outperforms state-of-the-art performance and the effectiveness of recommender. Junyan Jiang, Yanbing Liu 0007, Shujuan Chen |
IJCNN | 5 |
| 2020 | High Quality Candidate Generation and Sequential Graph Attention Network for Entity LinkingabstractEntity Linking (EL) is a task for mapping mentions in text to corresponding entities in knowledge base (KB). This task usually includes candidate generation (CG) and entity disambiguation (ED) stages. Recent EL systems based on neural network models have achieved good performance, but they still face two challenges: (i) Previous studies evaluate their models without considering the differences between candidate entities. In fact, the quality (gold recall in particular) of candidate sets has an effect on the EL results. So, how to promote the quality of candidates needs more attention. (ii) In order to utilize the topical coherence among the referred entities, many graph and sequence models are proposed for collective ED. However, graph-based models treat all candidate entities equally which may introduce much noise information. On the contrary, sequence models can only observe previous referred entities, ignoring the relevance between the current mention and its subsequent entities. To address the first problem, we propose a multi-strategy based CG method to generate high recall candidate sets. For the second problem, we design a Sequential Graph Attention Network (SeqGAT) which combines the advantages of graph and sequence methods. In our model, mentions are dealt with in a sequence manner. Given the current mention, SeqGAT dynamically encodes both its previous referred entities and subsequent ones, and assign different importance to these entities. In this way, it not only makes full use of the topical consistency, but also reduce noise interference. We conduct experiments on different types of datasets and compare our method with previous EL system on the open evaluation platform. The comparison results show that our model achieves significant improvements over the state-of-the-art methods. Zheng Fang 0002, Yanan Cao 0001, Zhenyu Zhang 0006, Yanbing Liu 0007, Shi Wang 0002 |
WWW | 5 |
| 2020 | Learning cross-modal correlations by exploring inter-word semantics and stacked co-attention
Jing Yu 0007, Weifeng Zhang 0002, Zengchang Qin, Yanbing Liu 0007, Yue Hu 0002 |
Pattern Recognit. Lett. | 5 |
| 2019 | PAAE: A Unified Framework for Predicting Anchor Links with Adversarial EmbeddingabstractThe goal of predicting anchor links is to align accounts from multiple networks by whether they are held by the same natural person. Network structure is the key information for predicting anchor links. Exploring the intrinsic attributes of the network structure is an important way to align anchor users across social networks. Existing methods use a representation learning approach to embed network vertices into low dimension vectors space. But these methods suffer from lack of additional constraints for enhancing the robustness of the embedding vectors when aligning anchor nodes across networks with large structural differences. To offer a robust method, we propose a novel adversarial representation learning approach to align users, called PAAE(predicting anchor links with adversarial embedding), which employs an adversarial regularization to capture the robust embedding vectors and maps anchor users with an alignment autoencoders. PAAE can solve both the network embedding problem and the user alignment problem simultaneously under a unified optimization framework. Through extensive experiments on real social network datasets, we demonstrate that PAAE significantly outperforms the state-of-the-art methods. Yanmin Shang, Zhezhou Kang, Yanan Cao 0001, Yangxi Li, Yanbing Liu 0007 |
ICME | 7 |
| 2019 | Combining Deep Neural Network with SVM to Identify Used in IOTabstractAlong with the development of science and technology, especially one of internet of things (IOT), products related to IOT have improved human life these days. In IOT-related utility products, it is impossible not to mention devices for smarts city, self-driving cars and especially for smart home. Devices for smarts home are usually controlled by voice. Therefore, voice processing technology is also in need of improvement. Speech processing enhancement that ensures safety in the process helps smart home devices bring about an evolutionary change. In the article, we mainly focus on human voice processing independently od the text. Particularly, we will integrate Convolutional network(CNN) and Suport Vector Machine(SVM) to create a Feature Building Machine. SVMs are often used in speech and image classification, which accordingly is a critical and swift data sorter. The article analyzes the advantages of the combination Deep Neural Network (DNN) and SVMs in speech recognition and is the foundation to develop devices for smart home. The results of the experiment, which was used in the standard Voxcelb database, demonstrate the superiority in sound recognition compared to traditional i-vector methods or other CNN methods. Nguyen Nang An, Nguyen Quang Thanh, Yanbing Liu 0007 |
IWCMC | 3 |
| 2019 | UAFA: Unsupervised Attribute-Friendship Attention Framework for User Representation
Yanmin Shang, Yaman Cao, Yanbing Liu 0007, Jianlong Tan |
KSEM (1) | 4 |
| 2019 | Joint Entity Linking with Deep Reinforcement LearningabstractEntity linking is the task of aligning mentions to corresponding entities in a given knowledge base. Previous studies have highlighted the necessity for entity linking systems to capture the global coherence. However, there are two common weaknesses in previous global models. First, most of them calculate the pairwise scores between all candidate entities and select the most relevant group of entities as the final result. In this process, the consistency among wrong entities as well as that among right ones are involved, which may introduce noise data and increase the model complexity. Second, the cues of previously disambiguated entities, which could contribute to the disambiguation of the subsequent mentions, are usually ignored by previous models. To address these problems, we convert the global linking into a sequence decision problem and propose a reinforcement learning model which makes decisions from a global perspective. Our model makes full use of the previous referred entities and explores the long-term influence of current selection on subsequent decisions. We conduct experiments on different types of datasets, the results show that our model outperforms state-of-the-art systems and has better generalization performance. Zheng Fang 0002, Yanan Cao 0001, Qian Li 0003, Zhenyu Zhang 0006, Yanbing Liu 0007 |
WWW | 6 |
| 2018 | Hierarchical Attention Networks for User Profile Inference in Social Media Systems
Zhezhou Kang, Yanan Cao 0001, Yanmin Shang, Yanbing Liu 0007, Li Guo 0001 |
ICANN (3) | 5 |
| 2018 | Reinforcement Learning for Joint Extraction of Entities and Relations
Wenpeng Liu, Yanan Cao 0001, Yanbing Liu 0007, Yue Hu 0002, Jianlong Tan |
ICANN (2) | 3 |
| 2018 | Attention-Based RNN Model for Joint Extraction of Intent and Word Slot Based on a Tagging Strategy
Zheng Fang 0002, Yanan Cao 0001, Yanbing Liu 0007, Xiaojun Chen 0004, Jianlong Tan |
ICANN (3) | 4 |
| 2018 | A Data-Deduplication-Based Matching Mechanism for URL FilteringabstractURL filtering plays an important role in various network security applications. URL filtering usually requires high matching performance, but the performance of the classical multiple string matching algorithms have been difficult to be significantly improved. In this article, we found that the online URLs to be filtered contain a large number of duplicate URLs. According to this observation, we propose a novel deduplication-based matching mechanism (DBM) for URL filtering. The DBM caches information of the duplicate URLs in a hash table to avoid duplicate URLs being repeatedly scanned by URL filtering system. The DBM can be used in conjunction with any multiple string matching algorithms. Experimental results show that when a multiple string matching algorithm used in conjunction with the DBM, the matching speed of the URL filtering system can be increased by 9\%-68\%. So DBM can significantly accelerate the speed of URL filtering system. Besides increasing speed of URL filtering system, DBM is a mechanism independent of the specific matching algorithm and can be easily used in other field. Yuhai Lu, Yanbing Liu 0007, Jianlong Tan |
ICC | 2 |
| 2018 | Sequence Generative Adversarial Network for Long Text SummarizationabstractIn this paper, we propose a new adversarial training framework for text summarization task. Although sequence-to-sequence models have achieved state-of-the-art performance in abstractive summarization, the training strategy (MLE) suffers from exposure bias in the inference stage. This discrepancy between training and inference makes generated summaries less coherent and accuracy, which is more prominent in summarizing long articles. To address this issue, we model abstractive summarization using Generative Adversarial Network (GAN), aiming to minimize the gap between generated summaries and the ground-truth ones. This framework consists of two models: a generator that generates summaries, a discriminator that evaluates generated summaries. Reinforcement learning (RL) strategy is used to guarantee the co-training of generator and discriminator. Besides, motivated by the nature of summarization task, we design a novel Triple-RNNs discriminator, and extend the off-the-shelf generator by appending encoder and decoder with attention mechanism. Experimental results showed that our model significantly outperforms the state-of-the-art models, especially on long text corpus. Yanan Cao 0001, Ruipeng Jia, Yanbing Liu 0007, Jianlong Tan |
ICTAI | 4 |
| 2018 | Fine-Grained Correlation Learning with Stacked Co-attention Networks for Cross-Modal Information Retrieval
Jing Yu 0007, Yanbing Liu 0007, Jianlong Tan, Li Guo 0001, Weifeng Zhang 0002 |
KSEM (1) | 3 |
| 2018 | A Sequence Transformation Model for Chinese Named Entity Recognition
Qingyue Wang, Yanjing Song, Yanan Cao 0001, Yanbing Liu 0007, Li Guo 0001 |
KSEM (1) | 5 |
| 2017 | Multiple Pattern Graph Correlations for Efficient Graph Pattern MatchingabstractGraph pattern matching has wide applications in social network analysis, such as identifying important communities, social roles and hidden behavior structures. Existing algorithms are mainly designed for single-pattern tasks, where patterns are processed sequentially and independently. However, many scenarios require multiple patterns to be processed as a batch, where single-pattern scheme will cost much redundant computation caused by matching similar sub-structures in the pattern set. Therefore, the sequential graph pattern matching scheme is not always the most efficient. This paper aims to propose a multiple pattern graph optimization algorithm for the subgraph isomorphism task. We comprehensively study the structural correlations among the multiple patterns and represent them by a compact tree-structured index. To support fast insertion and deletion of the pattern index, we present a dynamic updating algorithm to avoid index reconstruction from scratch. Based on the index, an efficient matching algorithm is proposed to answer multiple patterns in a heuristic scheduling order and avoid redundant computation. Extensive experiments on real and synthetic datasets prove that our solution is several times faster comparing with the state-of-the-art work. Jing Yu 0007, Yanbing Liu 0007, Yue Hu 0002 |
AICCSA | 3 |
| 2017 | Inferring Social Network User's Interest Based on Convolutional Neural Network
Yanan Cao 0001, Shi Wang 0002, Cong Cao 0001, Yanbing Liu 0007, Jianlong Tan |
ICONIP (5) | 5 |
| 2017 | Inferring User Profiles in Online Social Networks Based on Convolutional Neural Network
Yanan Cao 0001, Yanmin Shang, Yanbing Liu 0007, Jianlong Tan, Li Guo 0001 |
KSEM | 4 |
| 2014 | Delta-K 2-tree for Compact Representation of Web Graphs
Gang Xiong 0001, Yanbing Liu 0007, Ping Liu 0001, Li Guo 0001 |
APWeb | 3 |
| 2014 | A factor-searching-based multiple string matching algorithm for intrusion detectionabstractMultiple string matching plays a fundamental role in network intrusion detection systems. Automata-based multiple string matching algorithms like AC, SBDM and SBOM are widely used in practice, but the huge memory usage of automata prevents them from being applied to a large-scale pattern set. Meanwhile, poor cache locality of huge automata degrades the matching speed of algorithms. Here we propose a space-efficient multiple string matching algorithm BVM, which makes use of bit-vector and succinct hash table to replace the automata used in factor-searching-based algorithms. Space complexity of the proposed algorithm is O(rm2+ ΣpϵP|p|), that is more space-efficient than the classic automata-based algorithms. Experiments on datasets including Snort, ClamAV, URL blacklist and synthetic rules show that the proposed algorithm significantly reduces memory usage and still runs at a fast matching speed. Above all, BVM costs less than 0.75% of the memory usage of AC, and is capable of matching millions of patterns efficiently. Yanbing Liu 0007, Qingyun Liu 0001, Ping Liu 0001, Jianlong Tan, Li Guo 0001 |
ICC | 1 |
| 2012 | ClusterFA: a memory-efficient DFA structure for network intrusion detectionabstractNetwork intrusion detection systems (NIDS) plays an increasing important role in the field of network security. Current NIDS, such as Bro and Snort, mainly use signatures to represent and detect networking attacks. Traditionally the signatures are depicted by exact string patterns. However, new worms and viruses emerge endlessly in recent years. As a result, the scale of signatures increases sharply. Compared with exact strings, regular expressions have more powerful expressiveness, and are replacing exact strings gradually in state-of-the-art NIDS. Lei Jiang 0003, Jianlong Tan, Yanbing Liu 0007 |
AsiaCCS | 3 |
| 2011 | An efficient regular expressions compression algorithm from a new perspectiveabstractDeep packet inspection plays a increasingly important role in network security devices and applications, which use more regular expressions to depict patterns. DFA engine is usually used as a classical representation for regular expressions to perform pattern matching, because it only need O(1) time to process one input character. However, DFAs of regular expression sets require large amount of memory, which limits the practical application of regular expressions in high-speed networks. Some compression algorithms have been proposed to address this issue in recent literatures. In this paper, we reconsider this problem from a new perspective, namely observing the characteristic of transition distribution inside each state, which is different from previous algorithms that observe transition characteristic among states. Furthermore, we introduce a new compression algorithm which can reduce 95% memory usage of DFA stably without significant impact on matching speed. Moreover, our work is orthogonal to previous compression algorithms, such as D2FA, δFA. Our experiment results show that applying our work to them will have several times memory reduction, and matching speed of up to dozens of times comparing with original δFA in software implementation. Tingwen Liu, Yifu Yang, Yanbing Liu 0007, Li Guo 0001 |
INFOCOM | 3 |
| 2011 | Accelerating DFA Construction by Hierarchical MergingabstractRegular expression matching is widely used in many network applications to analyze suspicious traffic against predefined signatures, and to discover anomalous events. Deterministic Finite Automaton (DFA), which recognizes a set of regular expressions, is the basic data structure to scan input traffic byte by byte. Though DFA meets the requirement of real-time processing of network traffic, constructing a combined DFA for a set of regular expression signatures is very time-consuming, especially when the signature set is large. To attack this problem, we propose new strategies to accelerate DFA construction. The basic idea of our method is to construct the combined DFA by hierarchical merging of the DFAs of each single regular expression. Our method runs in $O(|Q| |\Sigma|\ln n)$ time, which is substantially superior to the time complexity $O(|Q| |\Sigma|(\overset{n}{\underset{i=1}{\sum}}|Q_i|)^2)$of classical subset construction algorithm\cite{Aho1986}. Experiment on real signatures from open-source systems, such as L7-filter, BRO and SNORT, demonstrates that our method performs 45 times faster than the subset construction algorithm on average. Yanbing Liu 0007, Li Guo 0001, Muyi Guo, Ping Liu 0001 |
ISPA | 1 |
| 2011 | Revisiting Multiple Pattern Matching Algorithms for Multi-Core Architecture
Guangming Tan, Ping Liu 0001, Dongbo Bu, Yanbing Liu 0007 |
J. Comput. Sci. Technol. | 4 |
| 2010 | Compressing Regular Expressions' DFA Table by Matrix Decomposition
Yanbing Liu 0007, Li Guo 0001, Ping Liu 0001, Jianlong Tan |
CIAA | 1 |
| 2009 | A Table Compression Method for Extended Aho-Corasick Automaton
Yanbing Liu 0007, Yifu Yang, Ping Liu 0001, Jianlong Tan |
CIAA | 1 |
| 2008 | Accelerating Multiple String Matching by Using Cache-Efficient StrategyabstractString matching plays a fundamental role in many network security applications such as NIDS, virus detection and information filtering. In this paper, we proposed cache-efficient methods to accelerate classical multiple string matching algorithms. We observed that most classical algorithms perform poorly as pattern set grows due to their high memory requirement and the poor cache behavior. Based on this observation, we proposed efficient methods employing cache-efficient strategies, i.e., to accelerate string matching by minimizing memory usage and maximizing cache locality. Experimental results on random datasets demonstrated that our new methods are substantially faster than classical methods. Jianlong Tan, Yanbing Liu 0007, Ping Liu 0001 |
WAIM | 2 |
| 2005 | A Partition-Based Efficient Algorithm for Large Scale Multiple-Strings Matching
Ping Liu 0001, Yanbing Liu 0007, Jianlong Tan |
SPIRE | 2 |