VLDB 2026 Research / reviewers in the wild / expert
Cong Cao 0001
dblp:342/1223
· DBLP profile ↗
40ranked-venue papers
0as first author
35since 2021 · last 2026
0000-0003-1881-1947ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 16 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 9 since 2021Human-computer interaction and ubiquitous computing · 7 · 7 since 2021Databases, data management, data science and information retrieval · 6 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Computer networks · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | OPERA: A Reinforcement Learning-Enhanced Orchestrated Planner-Executor Architecture for Reasoning-Oriented Multi-Hop RetrievalabstractRecent advances in large language models (LLMs) and dense retrievers have driven significant progress in retrieval-augmented generation (RAG). However, existing approaches face significant challenges in complex reasoning-oriented multi-hop retrieval tasks: 1) Ineffective reasoning-oriented planning: Prior methods struggle to generate robust multi-step plans for complex queries, as rule-based decomposers perform poorly on out-of-template questions. 2) Suboptimal reasoning-driven retrieval: Related methods employ limited query reformulation, leading to iterative retrieval loops that often fail to locate golden documents. 3) Insufficient reasoning-guided filtering: Prevailing methods lack the fine-grained reasoning to effectively filter salient information from noisy results, hindering utilization of retrieved knowledge. Fundamentally, these limitations all stem from the weak coupling between retrieval and reasoning in current RAG architectures. We introduce the Orchestrated Planner-Executor Reasoning Architecture (OPERA), a novel reasoning-driven retrieval framework. OPERA's Goal Planning Module (GPM) decomposes questions into sub-goals, which are executed by a Reason-Execute Module (REM) with specialized components for precise reasoning and effective retrieval. To train OPERA, we propose Multi-Agents Progressive Group Relative Policy Optimization (MAPGRPO), a novel variant of GRPO. Experiments on complex multi-hop benchmarks show OPERA's superior performance, validating both the MAPGRPO method and OPERA's design. Yanbing Liu 0007, Fangfang Yuan, Cong Cao 0001, Youbang Sun, Weizhuo Chen, Jianjun Li 0010, Zhiyuan Ma 0005 |
AAAI | 4 |
| 2026 | Two Streams, One Sarcasm: Orthogonal Expert Tuning for Holistic Multimodal Sarcasm UnderstandingabstractDiandian Guo, Cong Cao, Fangfang Yuan, Pin Xu, Cheng Hu, Zhicheng Zhang, Yu Liu, Yanbing Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Diandian Guo, Cong Cao 0001, Fangfang Yuan, Pin Xu, Yanbing Liu 0007 |
ACL (1) | 2 |
| 2026 | Trans-RAG: Query-Centric Vector Transformation for Secure Cross-Organizational Retrieval
Fangfang Yuan, Cong Cao 0001, Wenxuan Lu, Yanbing Liu 0007 |
DASFAA (4) | 5 |
| 2026 | FIRW: Frequency-injected Robust Watermarking for Latent Diffusion Models
Xiaokun Li, Fangfang Yuan, Cong Cao 0001, Majing Su, Yueshan Wang, Lei Jiang 0003, Yanbing Liu 0007 |
ICIC (2) | 3 |
| 2026 | MSD-NR: A Noise Rectification Method for Multimodal Sarcasm Detection
Xiangyu Tian, Diandian Guo, Cong Cao 0001, Yueshan Wang, Fangfang Yuan, Yanbing Liu 0007 |
ICIC | 3 |
| 2026 | Guided by LLM, Grounded in Targeted Vision: Dynamic Chain-of-Thought with Retrieval-Augmented In-Context Learning for Knowledge-Based Visual Question Answering
Yuling Yang, Weizhuo Chen, Cong Cao 0001, Fangfang Yuan, Yanbing Liu 0007 |
ICIC (19) | 3 |
| 2026 | TSS-RAG: Leveraging Teacher-Student Classroom Scenarios to Mitigate Hallucinations in Retrieval-Augmented Generation
Junqian Zhao, Tianhe Yu, Fangfang Yuan, Yanlong Zhou, Cong Cao 0001, Yanbing Liu 0007 |
ICIC | 6 |
| 2026 | MuVaC: A Variational Causal Framework for Multimodal Sarcasm Understanding in Dialogues
Diandian Guo, Fangfang Yuan, Cong Cao 0001, Xixun Lin, Chuan Zhou 0001, Hao Peng 0001, Yanan Cao 0006, Yanbing Liu 0007 |
WWW | 3 |
| 2025 | Context-Aware Heterogeneous Graph Interactive Learning for Document-Level Event Extraction
Fumin Guo, Fangfang Yuan, Cong Cao 0001 |
ADMA (2) | 5 |
| 2025 | Dialogues Aspect-based Sentiment Quadruple Extraction via Structural Entropy Minimization PartitioningabstractDialogues Aspect-based Sentiment Quadruple Extraction (DiaASQ) aims to extract all target-aspect-opinion-sentiment quadruples from a given multi-round, multi-participant dialogue. Existing methods typically learn word relations across entire dialogues, assuming a uniform distribution of sentiment elements. However, we find that dialogues often contain multiple semantically independent sub-dialogues without clear dependencies between them. Therefore, learning word relationships across the entire dialogue inevitably introduces additional noise into the extraction process. To address this, our method focuses on partitioning dialogues into semantically independent sub-dialogues. Achieving completeness while minimizing these sub-dialogues presents a significant challenge. Simply partitioning based on reply relationships is ineffective. Instead, we propose utilizing a structural entropy minimization algorithm to partition the dialogues. This approach aims to preserve relevant utterances while distinguishing irrelevant ones as much as possible. Furthermore, we introduce a two-step framework for quadruple extraction: first extracting individual sentiment elements at the utterance level, then matching quadruples at the sub-dialogue level. Extensive experiments demonstrate that our approach achieves state-of-the-art performance in DiaASQ with much lower computational costs. Cong Cao 0001, Hao Peng 0001, Zhifeng Hao 0004, Lei Jiang 0003, Kongjing Gu, Yanbing Liu 0007, Philip S. Yu |
CIKM | 2 |
| 2025 | Multi-View Incongruity Learning for Multimodal Sarcasm DetectionabstractMultimodal sarcasm detection (MSD) is essential for various downstream tasks. Existing MSD methods tend to rely on spurious correlations. These methods often mistakenly prioritize non-essential features yet still make correct predictions, demonstrating poor generalizability beyond training environments. Regarding this phenomenon, this paper undertakes several initiatives. Firstly, we identify two primary causes that lead to the reliance of spurious correlations. Secondly, we address these challenges by proposing a novel method that integrate Multimodal Incongruities via Contrastive Learning (MICL) for multimodal sarcasm detection. Specifically, we first leverage incongruity to drive multi-view learning from three views: token-patch, entity-object, and sentiment. Then, we introduce extensive data augmentation to mitigate the biased learning of the textual modality. Additionally, we construct a test set, SPMSD, which consists potential spurious correlations to evaluate the the model’s generalizability. Experimental results demonstrate the superiority of MICL on benchmark datasets, along with the analyses showcasing MICL’s advancement in mitigating the effect of spurious correlation. Diandian Guo, Cong Cao 0001, Fangfang Yuan, Yanbing Liu 0007, Guangjie Zeng, Hao Peng 0001, Philip S. Yu |
COLING | 2 |
| 2025 | Emotion Transfer with Enhanced Prototype for Unseen Emotion Recognition in ConversationabstractCurrent Emotion Recognition in Conversation (ERC) research follows a closed-domain assumption.However, there is no clear consensus on emotion classification in psychology, which presents a challenge for models when it comes to recognizing previously unseen emotions in real-world applications.To bridge this gap, we introduce the Unseen Emotion Recognition in Conversation (UERC) task for the first time and propose ProEmoTrans, a solid prototype-based emotion transfer framework.This prototype-based approach shows promise but still faces key challenges: First, implicit expressions complicate emotion definition, which we address by proposing an LLM-enhanced description approach.Second, utterance encoding in long conversations is difficult, which we tackle with a proposed parameter-free mechanism for efficient encoding and overfitting prevention.Finally, the Markovian flow nature of emotions is hard to transfer, which we address with an improved Attention Viterbi Decoding (AVD) method to transfer seen emotion transitions to unseen emotions.Extensive experiments on three datasets show that our method serves as a strong baseline for preliminary exploration in this new area. Cong Cao 0001, Hao Peng 0001, Guanlin Wu, Zhifeng Hao 0004, Lei Jiang 0003, Yanbing Liu 0007, Philip S. Yu |
EMNLP | 2 |
| 2025 | ReTD: Reconstruction-Based Traceability Detection for Generated ImagesabstractThe objective of generated image traceability is to accurately identify and locate the source models. In this paper, we propose ReTD (Reconstruction-Based Traceability Detection), a generalized model for generated image traceability detection. Firstly, we use VAE to reconstruct images which are compared with the original ones to extract numerical distinguishing features. Secondly, we use Vision Transformer to learn the fine-grained distribution features to realize the generated image traceability classification. Finally, we conduct traceability experiments using images generated by ten GAN and Diffusion models. The experimental results demonstrate that ReTD only training a unified classifier improves accuracy by 9.4% compared to the state-of-the-art method. The ReTD-related code is availble at https://github.com/chenweizhuo/ReTD. Weizhuo Chen, Fangfang Yuan, Cong Cao 0001, Dakui Wang, Yanbing Liu 0007 |
ICASSP | 3 |
| 2025 | Counterfactual-Augmented Representation Learning based Event PredictionabstractAccurate prediction of future events holds significant importance for decision-makers. Current methods learn representations of past events from the observable graph structure to predict whether a future event will occur. However, these methods overlook the counterfactual scenarios, thus missing essential factors that could trigger future events. In this paper, we propose the Predicting Events with Counterfactual Augmentation Framework (PECF) to address this limitation. This is achieved by investigating whether deviations from observed events (i.e., counterfactual events) can affect the occurrence of the target future event. Specifically, first, we learn the representations of events through the temporal event graphs. Then, we instantiate causal models to represent the causal relationships between events. Finally, we generate counterfactual events and enhance event prediction accuracy through counterfactual-based augmentation. Experimental results demonstrate that our method outperforms current state-of-the-art methods on benchmark datasets. The code is available at https://github.com/hucheng-IIE/PECF. Fangfang Yuan, Cong Cao 0001, Guangjie Zeng, Yanbing Liu 0007, Hao Peng 0001, Philip S. Yu |
ICME | 3 |
| 2025 | CASD: Counterfactual Augmentation for Social Bot Detection on TwitterabstractSocial bot detection has become increasingly important with the rise of social media platforms, as social bots could be used for malicious activities such as spreading misinformation. Recent advances mainly utilize Graph Neural Networks (GNNs) for bot detection. However, these detection methods may overlook the issue of social bots’ disguise. An advanced bot can highly mimic human users and homogenize their features with human accounts, ultimately weakening the effectiveness of social bot detectors. In this paper, we propose CASD, a counterfactual-based social bot detection method. Specifically, we design a subgraph generation module to enhance node feature learning by reducing interference from irrelevant global information in sparse networks. Then, we apply counterfactual reasoning to simulate the disguising of bots and identify their distinguishable features. Finally, we combine the counterfactual feature with its raw counterpart for robust bot detection. Experimental results on multiple datasets demonstrate that our method outperforms existing state-of-the-art (SOTA) methods. Our code is publicly available at1. Pin Xu, Fangfang Yuan, Yueshan Wang, Diandian Guo, Cong Cao 0001, Yanbing Liu 0007 |
ICME | 5 |
| 2025 | T-T: Table Transformer for Tagging-based Aspect Sentiment Triplet ExtractionabstractAspect sentiment triplet extraction (ASTE) aims to extract triplets composed of aspect terms, opinion terms, and sentiment polarities from given sentences. The table tagging method is a popular approach to addressing this task, which encodes a sentence into a 2-dimensional table, allowing for the tagging of relations between any two words. Previous efforts have focused on designing various downstream relation learning modules to better capture interactions between tokens in the table, revealing that a stronger capability in relation capture can lead to greater improvements in the model. Motivated by this, we attempt to directly utilize transformer layers as downstream relation learning modules. Due to the powerful semantic modeling capability of transformers, it is foreseeable that this will lead to excellent improvement. However, owing to the quadratic relation between the length of the table and the length of the input sentence sequence, using transformers directly faces two challenges: overly long table sequences and unfair local attention interaction. To address these challenges, we propose a novel Table-Transformer (T-T) for the tagging-based ASTE method. Specifically, we introduce a stripe attention mechanism with a loop-shift strategy to tackle these challenges. The former modifies the global attention mechanism to only attend to a 2-dimensional local attention window, while the latter facilitates interaction between different attention windows. Extensive and comprehensive experiments demonstrate that the T-T, as a downstream relation learning module, achieves state-of-the-art performance with lower computational costs. Chaodong Tong, Cong Cao 0001, Hao Peng 0001, Qian Li 0033, Guanlin Wu, Lei Jiang 0003, Yanbing Liu 0007, Philip S. Yu |
IJCAI | 3 |
| 2025 | See Better, Say Better: Vision-Augmented Decoding for Mitigating Hallucinations in Large Vision-Language Models
Xinyi Sun, Diandian Guo, Cong Cao 0001, Fangfang Yuan, Dakui Wang, Yanbing Liu 0007 |
NLPCC (1) | 3 |
| 2024 | Multi-Level Graph Convolutional Network for Document Information ExtractionabstractDocument information extraction aims to identify and extract entities and other essential information from documents. Its performance is significantly dependent on the learned relationship between the text and its corresponding layout, which is not fully exploited by existing methods. Therefore, we propose a multi-level graph convolutional network to improve the capability for modeling document information. Firstly, we integrate text embeddings with corresponding image and layout features as text features, which are used to construct a text-level graph to extract semantic relationships as word features through graph convolutions. We then construct a layout-level graph using region layout features as nodes, extracting structure relationships through further graph convolutions. This multi-level graph structure allows our model to fuse fine-grained word features with coarse-grained region features for effective sequence labeling. Comprehensive experiments on various datasets consistently demonstrate the effectiveness of our method, and the results show that our model outperforms the state-of-the-art methods with the aid of multi-level features. Fangfang Yuan, Dakui Wang, Cong Cao 0001, Yanbing Liu 0007 |
ICTAI | 6 |
| 2024 | Efficient One-Shot Pruning of Large Language Models with Low-Rank ApproximationabstractModel pruning, as an effective method for compressing large language models (LLMs), has recently attracted considerable attention in the field of natural language processing. However, existing LLM pruning methods have two main drawbacks: (1) Iterative pruning for LLMs with over a billion parameters requires retraining, which leads to significant pruning costs. (2) LLMs Pruning is formalized as a weight reconstruction problem that necessitates second-order information, incurring expensive computations. To address these issues, we propose a novel pruning method named Eplra: efficient one-shot pruning of large language models with low-rank approximation, which efficiently identifies sparse networks in LLMs. Specifically, we design a novel pruning metric based on input activations for the rapid one-shot compression of LLMs. We first incorporate input activations into the calculation of weight importance to promote precise pruning of low-priority weights. Then, we perform local weight comparisons across each output of linear layers to induce uniform sparsity. Next, we expand Eplra into semi-structured pruning patterns to accommodate various acceleration scenarios. Finally, we employ low-rank parametrized update matrices to fine-tune the pruned model, facilitating a swift recovery of model performance. Experimental results on various language benchmark datasets demonstrate that Eplra outperforms the state-of-the-art methods. Yangyan Xu, Cong Cao 0001, Fangfang Yuan, Rongxin Mi, Nannan Sun, Dakui Wang, Yanbing Liu 0007 |
SMC | 2 |
| 2024 | Enhancing GPT-3.5 for Knowledge-Based VQA with In-Context Prompt Learning and Image CaptioningabstractTraditional visual question answering (VQA) often falls short as merely relying on image information is insufficient to answer given questions. Therefore, Knowledge-Based Visual Question Answering (KB-VQA) has emerged. Typically, KB-VQA involves first retrieving knowledge from external knowledge bases, then using the retrieved knowledge in conjunction with the understanding of visual content for joint reasoning to predict answers. However, current models often suffer from weak visual perception capabilities when processing image information. Additionally, due to the incompleteness of external knowledge bases, retrieved knowledge may contain noise or even irrelevant information. Moreover, the re-embedding of knowledge text features during the model's reasoning process may deviate from the original meanings in the knowledge base. To address these challenges, we propose a method for Knowledge-Based Visual Question Answering (KB-VQA) using GPT-3.5, leveraging image captions and in-context prompts. We utilize an advanced captioning model to convert images into accurate textual representations, enhancing the large language model's understanding of visual information. Moreover, we eliminate the need for additional knowledge bases by directly employing GPT-3.5 as a knowledge base for knowledge retrieval and generate logically consistent text during inference to predict answers. Furthermore, we enhance GPT-3.5's question-answering capability for VQA through in-context prompt learning. Experiments on the public OK-VQA dataset demonstrate the superior performance of our model. Yuling Yang, Cong Cao 0001, Fangfang Yuan, Dakui Wang, Yanbing Liu 0007 |
SMC | 2 |
| 2023 | Confident Slot Iterative Learning for Multi-Domain Dialogue State Tracking
Qingyue Wang, Yanan Cao 0001, Piji Li, Yanhe Fu, Zheng Lin 0001, Cong Cao 0001, Shi Wang 0002, Li Guo 0001 |
CogSci | 6 |
| 2023 | NPGraph: An Efficient Graph Computing Model in NUMA-Based Persistent Memory Systems
Baoke Li, Cong Cao 0001, Fangfang Yuan, Yuling Yang, Majing Su, Yanbing Liu 0007, Jianhui Fu |
CollaborateCom (2) | 2 |
| 2023 | Few-shot Malicious Domain Detection on Heterogeneous Graph with Meta-learningabstractThe Domain Name System (DNS), one of the essential basic services on the Internet, is often abused by attackers to launch various cyber attacks, such as phishing and spamming. Researchers have proposed many machine learning-based and deep learning-based methods to detect malicious domains. However, these methods rely on a large-scale dataset with labeled samples for model training. The fact is that the labeled domain samples are limited in the real-world DNS dataset. In this paper, we propose a few-shot malicious domain detection model named MetaDom, which employs a meta-learning algorithm for model optimization. Specifically, We first model the DNS scenario as a heterogeneous graph to capture richer information by analysing the complex relations among domains, IP addresses and clients. Then, we learn the domain representations with a heterogeneous graph neural network on the DNS HG. Finally, considering that only few labeled data are available in the real-world DNS scenario, a meta-learning algorithm with knowledge distillation is introduced to optimize the model. Extensive experiments on the real DNS dataset show that MetaDom outperforms other state-of-the-art methods. Fangfang Yuan, Cong Cao 0001, Majing Su, Dakui Wang, Yanbing Liu 0007 |
CSCWD | 3 |
| 2023 | Curvature-Driven Knowledge Graph Embedding for Link PredictionabstractKnowledge Graph Embedding (KGE) aims to learn how to represent the low-dimensional vectors for entities and relations based on the observed triplets in knowledge graph. Most of the existing models use simple structural features, such as node degrees and directed edges, and pay little attention to advanced inherent information of structured knowledge. In this paper, we propose CD-GCN, a curvature-driven KGE method for link prediction. Specifically, we first apply Ricci curvature to knowledge graph. Then, we use curvature information to drive the state update, which aims to further exploit the graph-structured information. Finally, we use a ConvE scoring function to output the link prediction results. Through extensive experiments on public datasets FB15k-237 and WN18RR, CD-GCN has achieved state-of-the-art results compared with all baseline models. Diandian Guo, Majing Su, Cong Cao 0001, Fangfang Yuan, Yanbing Liu 0007, Jianhui Fu |
CSCWD | 3 |
| 2023 | MetaBERT: Collaborative Meta-Learning for Accelerating BERT InferenceabstractEarly exit methods are used to accelerate inference in pre-trained language models and maintain competitive performance on resource-constrained devices. However, existing methods for training early exit classifiers suffer from the problem of poor classifier representations in different layers, leading to difficulties in adapting to diverse natural language processing tasks. To address this issue, we propose MetaBERT: collaborative Meta-learning for accelerating BERT inference. The main goal of MetaBERT is to train early exit classifiers through collaborative meta-learning, in which case, few gradient updates can be quickly adapted to new tasks. Moreover, this novel meta-training approach produces good generalization performance, thus achieving an effective balance between the inference result and efficiency. Extensive experimental results show that our approach outperforms previous training methods by a large margin, and achieves state-of-the-art results compared to other competitive models. Yangyan Xu, Fangfang Yuan, Cong Cao 0001, Majing Su, Dakui Wang, Yanbing Liu 0007 |
CSCWD | 3 |
| 2023 | Malicious Domain Detection Based on Self-supervised HGNNs with Contrastive Learning
Zhiping Li, Fangfang Yuan, Cong Cao 0001, Majing Su, Yuhai Lu, Yanbing Liu 0007 |
ICANN (3) | 3 |
| 2023 | EDDVPL: A Web Attribute Extraction Method with Prompt Learning
Yuling Yang, Jiali Feng, Baoke Li, Fangfang Yuan, Cong Cao 0001, Yanbing Liu 0007 |
ICONIP (14) | 5 |
| 2023 | RegexClassifier: A GNN-Based Recognition Method for State-Explosive Regular ExpressionsabstractRegular expression (regex) matching technology has been widely used in various applications. For the sake of low time complexity and stable performance, Deterministic Finite Automaton (DFA) has become the first choice to perform fast regular expression matching. However, DFA has the state explosion problem, that is, the number of DFA states may increase exponentially while compiling some specific regexes to DFA. The huge memory consumption restricts its practical applications. A lot of works have addressed the DFA state explosion problem; however, none has met the requirements of fast recognition and small memory image. In this paper, we proposed RegexClassifier to recognize state-explosive regexes intelligently and efficiently. It firstly transforms regexes into Non-deterministic Finite Automatons(NFAs), then uses Graph Neural Network(GNN) models to classify NFAs in order to recognize regexes that may cause DFA state explosion. Experiments on typical rule sets show that the classification accuracy of the proposed model is up to 98%. Yuhai Lu, Fangfang Yuan, Cong Cao 0001, Yanbing Liu 0007 |
ISCC | 4 |
| 2023 | A Multi-granularity Similarity Enhanced Model for Implicit Event Argument Extraction
Yanhe Fu, Yi Liu 0067, Yanan Cao 0001, Yubing Ren, Qingyue Wang, Fang Fang 0009, Cong Cao 0001 |
NLPCC (2) | 7 |
| 2023 | Robust Malicious Domain Detection Against Adversarial Attacks on Heterogeneous GraphabstractDomain Name System (DNS) is a crucial infrastructure of the Internet, yet it is also a primary medium for disseminating illicit information. Researchers have proposed numerous methods to detect malicious domains, among which heterogeneous graph (HG) based models have demonstrated good performance. However, their success may also motivate attackers to defeat HG based models in order to evade detection. In this paper, we propose a novel malicious domain detection model named RoDom, which is robust against adversarial attacks on HG. Firstly, we introduce different perturbations to construct multiple attacked graphs, which are designed to simulate different types of adversarial attacks on the HG. Secondly, we design a discriminator to perform robust representation learning on the HG by discriminating the original graph from attacked graphs. Finally, we introduce a classification selector to further improve the model's robustness by automatically combining domain representations of multiple HGs for domain classification. The experimental results show that RoDom out-performs other state-of-the-art methods and exhibits stronger robustness against adversarial attacks on the HG. Zhiping Li, Fangfang Yuan, Dakui Wang, Cong Cao 0001, Yanbing Liu 0007 |
SMC | 6 |
| 2022 | Heterogeneous Graph Attention Network for Malicious Domain Detection
Zhiping Li, Fangfang Yuan, Yanbing Liu 0007, Cong Cao 0001, Fang Fang 0009, Jianlong Tan |
ICANN (2) | 4 |
| 2022 | DOM2R-Graph: A Web Attribute Extraction Architecture with Relation-Aware Heterogeneous Graph Transformer
Jiali Feng, Cong Cao 0001, Fangfang Yuan, Zhiping Li, Yanbing Liu 0007, Jianlong Tan |
ICONIP (1) | 2 |
| 2022 | Malicious Domain Detection with Heterogeneous Graph Propagation Network
Fangfang Yuan, Yanbing Liu 0007, Cong Cao 0001, Jianlong Tan |
WASA (1) | 4 |
| 2021 | HSAN: A Hierarchical Self-Attention Network for Multi-Turn Dialogue GenerationabstractIn the multi-turn dialogue system, response generation is not only related to the sentences in context but also relies on the words in each utterance. Although there are lots of methods that pay attention to model words and utterances, there still exist problems such as tending to generate common responses. In this paper, we propose a hierarchical self-attention network, named HSAN, which attends to the important words and utterances in context simultaneously. Firstly, we use the hierarchical encoder to update the word and utterance representations with their position information respectively. Secondly, the response representations are updated by the mask self-attention module in the decoder. Finally, the relevance between utterances and response is computed by another self-attention module and used for the next response decoding process. In terms of automatic metrics and human judgements, experimental results show that HSAN significantly outperforms all baselines on two common public datasets. Yawei Kong, Lu Zhang 0084, Can Ma, Cong Cao 0001 |
ICASSP | 4 |
| 2021 | Knowledge-Based Diverse Feature Transformation for Few-Shot Relation Classification
Yubao Tang, Zhezhou Li, Cong Cao 0001, Fang Fang 0009, Yanan Cao 0001, Yanbing Liu 0007, Jianhui Fu |
KSEM | 3 |
| 2020 | Neural Extractive Summarization with Hierarchical Attentive Heterogeneous Graph NetworkabstractSentence-level extractive text summarization is substantially a node classification task of network mining, adhering to the informative components and concise representations.There are lots of redundant phrases between extracted sentences, but it is difficult to model them exactly by the general supervised methods.Previous sentence encoders, especially BERT, specialize in modeling the relationship between source sentences.While, they have no ability to consider the overlaps of the target selected summary, and there are inherent dependencies among target labels of sentences.In this paper, we propose HAHSum (as shorthand for Hierarchical Attentive Heterogeneous Graph for Text Summarization), which well models different levels of information, including words and sentences, and spotlights redundancy dependencies between sentences.Our approach iteratively refines the sentence representations with redundancy-aware graph and delivers the label dependencies by message passing.Experiments on large scale benchmark corpus (CNN/DM, NYT, and NEWSROOM) demonstrate that HAHSum yields ground-breaking performance and outperforms previous extractive summarizers. Ruipeng Jia, Yanan Cao 0001, Hengzhu Tang, Fang Fang 0009, Cong Cao 0001, Shi Wang 0002 |
EMNLP (1) | 5 |
| 2017 | Inferring Social Network User's Interest Based on Convolutional Neural Network
Yanan Cao 0001, Shi Wang 0002, Cong Cao 0001, Yanbing Liu 0007, Jianlong Tan |
ICONIP (5) | 4 |
| 2015 | A Chinese Framework of Semantic Taxonomy and Description: Preliminary Experimental Evaluation Using Web Information ExtractionabstractThe Chinese Framework of Semantic Taxonomy and Description (FSTD) is a linguistic resource that stores lexical and predicate-argument semantics about events or states in Chinese text, developed with the application of knowledge acquisition from Chinese text in mind. In this paper we build a web information extraction system, called NkiExtractor, to evaluate FSTD experimentally. We use two metrics: grammar coverage measures whether there is a semantic category of FSTD that corresponds to an event description in text, and extraction precision measures whether the correct predicate-argument structure can be extracted from text. Experimental results show that FSTD is a fairly comprehensive and effective resource for knowledge acquisition. We also discuss future work for expanding FSTD and improving extraction precision of NkiExtractor. Liangjun Zang, Weimin Wang 0002, Fang Fang 0009, Cong Cao 0001, Cun-gen Cao 0001 |
KSEM | 5 |
| 2013 | A Survey of Commonsense Knowledge Acquisition
Liangjun Zang, Cong Cao 0001, Yanan Cao 0001, Yuming Wu, Cun-gen Cao 0001 |
J. Comput. Sci. Technol. | 2 |
| 2011 | An efficient approach to representing and mining knowledge from Qing court medical records
Weimin Wang 0002, Jingchun Zhang, Cong Cao 0001, Yue Liu 0035, Keji Chen |
Frontiers Comput. Sci. China | 3 |