EDBT 2026 Demo / reviewers in the wild / expert
Fang Fang 0009
dblp:74/3719-9
· DBLP profile ↗
32ranked-venue papers
1as first author
25since 2021 · last 2026
0009-0005-2907-2643ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 26 · 1 first-author · 20 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MetaGDPO: Alleviating Catastrophic Forgetting with Metacognitive Knowledge Through Group Direct Preference OptimizationabstractLarge Language Models demonstrate strong reasoning capabilities, which can be effectively compressed into smaller models. However, existing datasets and fine-tuning approaches still face challenges that lead to catastrophic forgetting, particularly for models smaller than 8B. First, most datasets typically ignore the relationship between training data knowledge and the model's inherent abilities, making it difficult to preserve prior knowledge. Second, conventional training objectives often fail to constrain inherent knowledge preservation, which can result in forgetting of previously learned skills. To address these issues, we propose a comprehensive solution that alleviates catastrophic forgetting from both the data and fine-tuning approach perspectives. On the data side, we construct a dataset of 5K instances that covers multiple reasoning tasks and incorporates metacognitive knowledge, making it more tolerant and effective for distillation into smaller models. We annotate the metacognitive knowledge required to solve each question and filter the data based on task knowledge and the model's inherent skills. On the training side, we introduce GDPO (Group Direction Preference Optimization), which is better suited for resource-limited scenarios and can efficiently approximate the performance of GRPO. Guided by the large model and by implicitly constraining the optimization path through a reference model, GDPO enables more effective knowledge transfer from the large model and constrains excessive parameter drift. Extensive experiments demonstrate that our approach significantly alleviates catastrophic forgetting and improves reasoning performance on smaller models. Lanxue Zhang, Yuqiang Xie, Fang Fang 0009, Fanglong Dong, Rui Liu 0032, Yanan Cao 0001 |
AAAI | 3 |
| 2026 | Do LLMs Know Tool Irrelevance? Demystifying Structural Alignment Bias in Tool InvocationsabstractLarge language models (LLMs) have demonstrated impressive capabilities in utilizing external tools.In practice, however, LLMs are often exposed to tools that are irrelevant to the user's query, in which case the desired behavior is to refrain from invocations.In this work, we identify a widespread yet overlooked mechanistic flaw in tool refusal, which we term structural alignment bias: Even when a tool fails to serve the user's goal, LLMs still tend to invoke it whenever query attributes can be validly assigned to tool parameters.To systematically study this bias, we introduce SABEval, a new dataset that decouples structural alignment from semantic relevance.Our analysis shows that structural alignment bias induces severe tool-invocation errors in LLMs, yet remains largely unaccounted for in existing evaluations.To investigate the internal mechanisms underlying this bias, we propose Contrastive Attention Attribution, which reveals two competing pathways for semantic checking and structural matching.The relative strength of these pathways drives LLMs' tool invocation decisions.Based on these findings, we further introduce a rebalancing strategy that effectively mitigates structural alignment bias, as demonstrated by extensive experiments, without degrading general tool-use capabilities.Our code and dataset are available at https://github. com/along-l/irrelevant-tool. Can you tell me if 'Spider-Man: Miles Morales' is on sale on the Canadian PlayStation Store?"name": "check_nintendo_switch_game_on_sale", "description": "Checks if a particular game is currently on sale on the Nintendo Switch Store.","parameters": { "game_title": {"type": "string", "description": "The title of the video game."},"region": {"type": "string", "description": "The region to check for the Nintendo Switch Store."}} Xixun Lin, Ge Zhang 0002, Fang Fang 0009, Yanan Cao 0001 |
ACL (1) | 5 |
| 2026 | EA-Agent: A Structured Multi-Step Reasoning Agent for Entity AlignmentabstractEntity alignment (EA) aims to identify entities across different knowledge graphs (KGs) that refer to the same real-world object and plays a critical role in knowledge fusion and integration.Traditional EA methods mainly rely on knowledge representation learning, but their performance is often limited under noisy or sparsely supervised scenarios.Recently, large language models (LLMs) have been introduced to EA and achieved notable improvements by leveraging rich semantic knowledge.However, existing LLM-based EA approaches typically treat LLMs as black-box decision makers, resulting in limited interpretability, and the direct use of large-scale triples substantially increases inference cost.To address these challenges, we propose EA-Agent, a reasoningdriven agent for EA.EA-Agent formulates EA as a structured reasoning process with multistep planning and execution, enabling interpretable alignment decisions.Within this process, it introduces attribute and relation triple selectors to filter redundant triples before feeding them into the LLM, effectively addressing efficiency challenges.Experimental results on three benchmark datasets demonstrate that EA-Agent consistently outperforms existing EA methods and achieves state-of-the-art performance.The source code is available at https: //github.com/YXNan0110/EA-Agent. Yixuan Nan, Xixun Lin, Yanmin Shang, Ge Zhang 0002, Zheng Fang 0002, Fang Fang 0009, Yanan Cao 0001 |
ACL (1) | 6 |
| 2026 | Exons-Detect: Identifying and Amplifying Exonic Tokens via Hidden-State Discrepancy for Robust AI-Generated Text DetectionabstractThe rapid advancement of large language models has increasingly blurred the boundary between human-written and AI-generated text, raising societal risks such as misinformation dissemination, authorship ambiguity, and threats to intellectual property rights.These concerns highlight the urgent need for effective and reliable detection methods.While existing training-free approaches often achieve strong performance by aggregating token-level signals into a global score, they typically assume uniform token contributions, making them less robust under short sequences or localized token modifications.To address these limitations, we propose Exons-Detect, a training-free method for AI-generated text detection based on an exon-aware token reweighting perspective.Exons-Detect identifies and amplifies informative exonic tokens by measuring hiddenstate discrepancy under a dual-model setting, and computes an interpretable translation score from the resulting importance-weighted token sequence.Empirical evaluations demonstrate that Exons-Detect achieves state-of-the-art detection performance and exhibits strong robustness to adversarial attacks and varying input lengths.In particular, it attains a 2.2% relative improvement in average AUROC over the strongest prior baseline on DetectRL.Code and data are available at https://github.com/ Xiaoweizhu57/Exons-Detect. Yubing Ren, Fang Fang 0009, Shi Wang 0002, Yanan Cao 0001, Li Guo 0001 |
ACL (1) | 3 |
| 2025 | PIG: Privacy Jailbreak Attack on LLMs via Gradient-based Iterative In-Context OptimizationabstractLarge Language Models (LLMs) excel in various domains but pose inherent privacy risks.Existing methods to evaluate privacy leakage in LLMs often use memorized prefixes or simple instructions to extract data, both of which well-alignment models can easily block.Meanwhile, Jailbreak attacks bypass LLM safety mechanisms to generate harmful content, but their role in privacy scenarios remains underexplored.In this paper, we examine the effectiveness of jailbreak attacks in extracting sensitive information, bridging privacy leakage and jailbreak attacks in LLMs.Moreover, we propose PIG, a novel framework targeting Personally Identifiable Information (PII) and addressing the limitations of current jailbreak methods.Specifically, PIG identifies PII entities and their types in privacy queries, uses in-context learning to build a privacy context, and iteratively updates it with three gradient-based strategies to elicit target PII.We evaluate PIG and existing jailbreak methods using two privacy-related datasets.Experiments on four white-box and two blackbox LLMs show that PIG outperforms baseline methods and achieves state-of-the-art (SoTA) results.The results underscore significant privacy risks in LLMs, emphasizing the need for stronger safeguards. Yidan Wang 0001, Yanan Cao 0001, Yubing Ren, Fang Fang 0009, Zheng Lin 0001, Binxing Fang |
ACL (1) | 4 |
| 2025 | Dynamic Evaluation with Cognitive Reasoning for Multi-turn Safety of Large Language ModelsabstractThe rapid advancement of Large Language Models (LLMs) poses significant challenges for safety evaluation. Current static datasets struggle to identify emerging vulnerabilities due to three limitations: (1) they risk being exposed in model training data, leading to evaluation bias; (2) their limited prompt diversity fails to capture real-world application scenarios; (3) they are limited to provide human-like multi-turn interactions. To address these limitations, we propose a dynamic evaluation framework, CogSafe, for comprehensive and automated multi-turn safety assessment of LLMs. We introduce CogSafe based on cognitive theories to simulate the real chatting process. To enhance assessment diversity, we introduce scenario simulation and strategy decision to guide the dynamic generation, enabling coverage of application situations. Furthermore, we incorporate the cognitive process to simulate multi-turn dialogues that reflect the cognitive dynamics of real-world interactions. Extensive experiments demonstrate the scalability and effectiveness of our framework, which has been applied to evaluate the safety of widely used LLMs. Lanxue Zhang, Yanan Cao 0001, Yuqiang Xie, Fang Fang 0009, Yangxi Li |
ACL (1) | 4 |
| 2025 | Reliably Bounding False Positives: A Zero-Shot Machine-Generated Text Detection Framework via Multiscaled Conformal PredictionabstractThe rapid advancement of large language models has raised significant concerns regarding their potential misuse by malicious actors.As a result, developing effective detectors to mitigate these risks has become a critical priority.However, most existing detection methods focus excessively on detection accuracy, often neglecting the societal risks posed by high false positive rates (FPRs).This paper addresses this issue by leveraging Conformal Prediction (CP), which effectively constrains the upper bound of FPRs.While directly applying CP constraints FPRs, it also leads to a significant reduction in detection performance.To overcome this trade-off, this paper proposes a Zero-Shot Machine-Generated Text Detection Framework via Multiscaled Conformal Prediction (MCP), which both enforces the FPR constraint and improves detection performance.This paper also introduces RealDet, a high-quality dataset that spans a wide range of domains, ensuring realistic calibration and enabling superior detection performance when combined with MCP.Empirical evaluations demonstrate that MCP effectively constrains FPRs, significantly enhances detection performance, and increases robustness against adversarial attacks across multiple detectors and datasets. Yubing Ren, Yanan Cao 0001, Xixun Lin, Fang Fang 0009, Yangxi Li |
ACL (1) | 5 |
| 2025 | ToP: a Structured Pathway for Document-level Information Extraction with Large Language ModelsabstractIn Natural Language Processing (NLP), Document-level Information Extraction (DocIE) poses a significant challenge, requiring analyzing and synthesizing information across extensive contexts. This complexity extends beyond the scale of data, delving into the nuanced interplay of context, semantics, and text elements. Recent advancements in Large Language Models (LLMs) have opened new prospects for DocIE. However, their application is limited by the models’ struggles with extended contexts and complex narrative structures, a gap highlighted when comparing the efficacy of direct LLM prompting against Chain-of-Thought (CoT) prompting. Addressing these challenges, we introduce the Tree of Plans (Top) framework, a novel approach that mirrors human decision-making by engaging in slow, deliberate reasoning. Top decomposes the DocIE task into three structured stages: (1) Task-specific Seed Plan Decomposition, (2) LLM-driven Plan Sampling and Voting, and (3) Instruction-prompted Plan Execution. Each stage leverages zero-shot prompting of an LLM, ensuring coherent progression and comprehensive document-level analysis. We evaluate Top on four challenging DocIE datasets, which cover document-level relation extraction and document-level event extraction tasks, both requiring complex reasoning over the long context. Pengfei Yin, Yubing Ren, Fang Fang 0009, Boxiang Hu |
IJCNN | 3 |
| 2025 | DNA-DetectLLM: Unveiling AI-Generated Text via a DNA-Inspired Mutation-Repair ParadigmabstractThe rapid advancement of large language models (LLMs) has blurred the line between AI-generated and human-written text. This progress brings societal risks such as misinformation, authorship ambiguity, and intellectual property concerns, highlighting the urgent need for reliable AI-generated text detection methods. However, recent advances in generative language modeling have resulted in significant overlap between the feature distributions of human-written and AI-generated text, blurring classification boundaries and making accurate detection increasingly challenging. To address the above challenges, we propose a DNA-inspired perspective, leveraging a repair-based process to directly and interpretably capture the intrinsic differences between human-written and AI-generated text. Building on this perspective, we introduce **DNA-DetectLLM**, a zero-shot detection method for distinguishing AI-generated and human-written text. The method constructs an ideal AI-generated sequence for each input, iteratively repairs non-optimal tokens, and quantifies the cumulative repair effort as an interpretable detection signal. Empirical evaluations demonstrate that our method achieves state-of-the-art detection performance and exhibits strong robustness against various adversarial attacks and input lengths. Specifically, DNA-DetectLLM achieves relative improvements of **5.55\%** in AUROC and **2.08\%** in F1 score across multiple public benchmark datasets. Code and data are available at https://github.com/Xiaoweizhu57/DNA-DetectLLM. Yubing Ren, Fang Fang 0009, Qingfeng Tan, Shi Wang 0002, Yanan Cao 0001 |
NeurIPS | 3 |
| 2025 | Bridging the Gap: Aligning Language Model Generation with Structured Information Extraction via Controllable State TransitionabstractLarge language models (LLMs) achieve superior performance in generative tasks. However, due to the natural gap between language model generation and structured information extraction in three dimensions: task type, output format, and modeling granularity, they often fall short in structured information extraction, a crucial capability for effective data utilization on the web. In this paper, we define the generation process of the language model as the controllable state transition, aligning the generation and extraction processes to ensure the integrity of the output structure and adapt to the goals of the information extraction task. Furthermore, we propose the Structure2Text decider to help the language model understand the fine-grained extraction information, which converts the structured output into natural language and makes state decisions, thereby focusing on the task-specific information kernels, and alleviating language model hallucinations and incorrect content generation. We conduct extensive experiments and detailed analyses on myriad information extraction tasks, including named entity recognition, relation extraction, and event argument extraction. Our method not only achieves significant performance improvements but also considerably enhances the model's capability to generate precise and relevant content, making the extracted content easy to parse. Hao Li 0156, Yubing Ren, Yanan Cao 0001, Fang Fang 0009, Zheng Lin 0001, Shi Wang 0002 |
WWW | 5 |
| 2024 | DEIE: Benchmarking Document-level Event Information Extraction with a Large-scale Chinese News DatasetabstractA text corpus centered on events is foundational to research concerning the detection, representation, reasoning, and harnessing of online events. The majority of current event-based datasets mainly target sentence-level tasks, thus to advance event-related research spanning from sentence to document level, this paper introduces DEIE, a unified large-scale document-level event information extraction dataset with over 56,000+ events and 242,000+ arguments. Three key features stand out: large-scale manual annotation (20,000 documents), comprehensive unified annotation (encompassing event trigger/argument, summary, and relation at once), and emergency events annotation (covering 19 emergency types). Notably, our experiments reveal that current event-related models struggle with DEIE, signaling a pressing need for more advanced event-related research in the future. Yubing Ren, Yanan Cao 0001, Hao Li 0156, Zixuan ZM Ma, Fang Fang 0009, Ping Guo 0002 |
LREC/COLING | 6 |
| 2024 | Sorting, Reasoning, and Extraction: An Easy-to-Hard Reasoning Framework for Document-Level Event Argument ExtractionabstractDocument-level event argument extraction is a crucial task to help understand event information. Existing methods mostly ignore the different extraction difficulties of arguments, and the lack of task planning significantly affects the extraction and reasoning abilities of the model. In this paper, we innovatively analyze the difficulty of arguments and propose a novel framework for reasoning from easy to hard, aiming to use the information of simple arguments to help the extraction of difficult arguments in a human-like way. Specifically, our framework consists of three core modules: sorting, reasoning, and extraction. The sorting module first sorts the argument roles according to the current context and plans the reasoning path from easy to hard. Then, the reasoning module performs information reasoning based on the reasoning path to help capture the information of difficult arguments. Finally, the extraction module utilizes the reasoning information to complete argument extraction. Experimental results on the RAMS and WikiEvents datasets show the great advantages of our proposed approach. In particular, we obtain new state-of-the-art (SOTA) performance in multiple scenarios. Hao Li 0156, Yanan Cao 0001, Yubing Ren, Fang Fang 0009, Lanxue Zhang, Shi Wang 0002 |
ICASSP | 4 |
| 2023 | Towards Better Entity Linking with Multi-View Enhanced DistillationabstractYi Liu, Yuan Tian, Jianxun Lian, Xinlong Wang, Yanan Cao, Fang Fang, Wen Zhang, Haizhen Huang, Weiwei Deng, Qi Zhang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Yi Liu 0067, Jianxun Lian, Yanan Cao 0001, Fang Fang 0009, Haizhen Huang, Qi Zhang 0066 |
ACL (1) | 6 |
| 2023 | Retrieve-and-Sample: Document-level Event Argument Extraction via Hybrid Retrieval AugmentationabstractRecent studies have shown the effectiveness of retrieval augmentation in many generative NLP tasks.These retrieval-augmented methods allow models to explicitly acquire prior external knowledge in a non-parametric manner and regard the retrieved reference instances as cues to augment text generation.These methods use similarity-based retrieval, which is based on a simple hypothesis: the more the retrieved demonstration resembles the original input, the more likely the demonstration label resembles the input label.However, due to the complexity of event labels and sparsity of event arguments, this hypothesis does not always hold in document-level EAE.This raises an interesting question: How do we design the retrieval strategy for document-level EAE?We investigate various retrieval settings from the input and label distribution views in this paper.We further augment document-level EAE with pseudo demonstrations sampled from event semantic regions that can cover adequate alternatives in the same context and event schema.Through extensive experiments on RAMS and WikiEvents, we demonstrate the validity of our newly introduced retrieval-augmented methods and analyze why they work. Yubing Ren, Yanan Cao 0001, Ping Guo 0002, Fang Fang 0009, Zheng Lin 0001 |
ACL (1) | 4 |
| 2023 | Seri: Sketching-Reasoning-Integrating Progressive Workflow for Empathetic Response GenerationabstractEmpathy is a key ability for a human-like dialogue system. Inspired by social psychology, empathy includes both affective and cognitive aspects. Previous works on this topic have merely focused on recognizing emotions or modeling cognition with commonsense knowledge. Nevertheless, the generated results of these works still have a big gap with human-like empathetic responses. In this paper, we propose Seri, a SkEtching-Reasoning-Integrating framework for empathetic response generation. In particular, we define an empathy planner to capture and reason about multi-source information that considers cognition and affection. Further, we introduce a dynamic integrator module that allows the model dynamically select the appropriate information to generate empathetic responses. Experimental results on EmpatheticDialogue show that our method outperforms competitive baselines and generates responses with higher diversity and cognitive empathy levels. Guanqun Bi, Yanan Cao 0001, Piji Li, Yuqiang Xie, Fang Fang 0009, Zheng Lin 0001 |
ICASSP | 5 |
| 2023 | A Multi-granularity Similarity Enhanced Model for Implicit Event Argument Extraction
Yanhe Fu, Yi Liu 0067, Yanan Cao 0001, Yubing Ren, Qingyue Wang, Fang Fang 0009, Cong Cao 0001 |
NLPCC (2) | 6 |
| 2022 | How Does Knowledge Graph Embedding Extrapolate to Unseen Data: A Semantic Evidence ViewabstractKnowledge Graph Embedding (KGE) aims to learn representations for entities and relations. Most KGE models have gained great success, especially on extrapolation scenarios. Specifically, given an unseen triple (h, r, t), a trained model can still correctly predict t from (h, r, ?), or h from (?, r, t), such extrapolation ability is impressive. However, most existing KGE works focus on the design of delicate triple modeling function, which mainly tells us how to measure the plausibility of observed triples, but offers limited explanation of why the methods can extrapolate to unseen data, and what are the important factors to help KGE extrapolate. Therefore in this work, we attempt to study the KGE extrapolation of two problems: 1. How does KGE extrapolate to unseen data? 2. How to design the KGE model with better extrapolation ability? For the problem 1, we first discuss the impact factors for extrapolation and from relation, entity and triple level respectively, propose three Semantic Evidences (SEs), which can be observed from train set and provide important semantic information for extrapolation. Then we verify the effectiveness of SEs through extensive experiments on several typical KGE methods. For the problem 2, to make better use of the three levels of SE, we propose a novel GNN-based KGE model, called Semantic Evidence aware Graph Neural Network (SE-GNN). In SE-GNN, each level of SE is modeled explicitly by the corresponding neighbor pattern, and merged sufficiently by the multi-layer aggregation, which contributes to obtaining more extrapolative knowledge representation. Finally, through extensive experiments on FB15k-237 and WN18RR datasets, we show that SE-GNN achieves state-of-the-art performance on Knowledge Graph Completion task and performs a better extrapolation ability. Our code is available at https://github.com/renli1024/SE-GNN. Yanan Cao 0001, Qiannan Zhu, Guanqun Bi, Fang Fang 0009, Yi Liu 0067, Qian Li 0003 |
AAAI | 5 |
| 2022 | CLIO: Role-interactive Multi-event Head Attention Network for Document-level Event ExtractionabstractTransforming the large amounts of unstructured text on the Internet into structured event knowledge is a critical, yet unsolved goal of NLP, especially when addressing document-level text. Existing methods struggle in Document-level Event Extraction (DEE) due to its two intrinsic challenges: (a) Nested arguments, which means one argument is the sub-string of another one. (b) Multiple events, which indicates we should identify multiple events and assemble the arguments for them. In this paper, we propose a role-interactive multi-event head attention network (CLIO) to solve these two challenges jointly. The key idea is to map different events to multiple subspaces (i.e. multi-event head). In each event subspace, we draw the semantic representation of each role closer to its corresponding arguments, then we determine whether the current event exists. To further optimize event representation, we propose an event representation enhancing strategy to regularize pre-trained embedding space to be more isotropic. Our experiments on two widely used DEE datasets show that CLIO achieves consistent improvements over previous methods. Yubing Ren, Yanan Cao 0001, Fang Fang 0009, Ping Guo 0002, Zheng Lin 0001, Yi Liu 0067 |
COLING | 3 |
| 2022 | Heterogeneous Graph Attention Network for Malicious Domain Detection
Zhiping Li, Fangfang Yuan, Yanbing Liu 0007, Cong Cao 0001, Fang Fang 0009, Jianlong Tan |
ICANN (2) | 5 |
| 2021 | Flexible Non-Autoregressive Extractive Summarization with Threshold: How to Extract a Non-Fixed Number of Summary SentencesabstractSentence-level extractive summarization is a fundamental yet challenging task, and recent powerful approaches prefer to pick sentences sorted by the predicted probabilities until the length limit is reached, a.k.a. ``Top-K Strategy''. This length limit is fixed based on the validation set, resulting in the lack of flexibility. In this work, we propose a more flexible and accurate non-autoregressive method for single document extractive summarization, extracting a non-fixed number of summary sentences without the sorting step. We call our approach ThresSum as it picks sentences simultaneously and individually from the source document when the predicted probabilities exceed a threshold. During training, the model enhances sentence representation through iterative refinement and the intermediate latent variables receive some weak supervision with soft labels, which are generated progressively by adjusting the temperature with a knowledge distillation algorithm. Specifically, the temperature is initialized with high value and drops along with the iteration until a temperature of 1. Experimental results on CNN/DM and NYT datasets have demonstrated the effectiveness of ThresSum, which significantly outperforms BERTSUMEXT with a substantial improvement of 0.74 ROUGE-1 score on CNN/DM. Our source code will be available on Github. Ruipeng Jia, Yanan Cao 0001, Haichao Shi, Fang Fang 0009, Pengfei Yin, Shi Wang 0002 |
AAAI | 4 |
| 2021 | Deep Differential Amplifier for Extractive SummarizationabstractRuipeng Jia, Yanan Cao, Fang Fang, Yuchen Zhou, Zheng Fang, Yanbing Liu, Shi Wang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Ruipeng Jia, Yanan Cao 0001, Fang Fang 0009, Zheng Fang 0002, Yanbing Liu 0007, Shi Wang 0002 |
ACL/IJCNLP (1) | 3 |
| 2021 | TEBNER: Domain Specific Named Entity Recognition with Type Expanded Boundary-aware NetworkabstractTo alleviate label scarcity in Named Entity Recognition (NER) task, distantly supervised NER methods are widely applied to automatically label data and identify entities.Although the human effort is reduced, the generated incomplete and noisy annotations pose new challenges for learning effective neural models.In this paper, we propose a novel dictionary extension method which extracts new entities through the type expanded model.Moreover, we design a multi-granularity boundaryaware network which detects entity boundaries from both local and global perspectives.We conduct experiments on different types of datasets, the results show that our model outperforms previous state-of-the-art distantly supervised systems and even surpasses the supervised models. Zheng Fang 0002, Yanan Cao 0001, Tai Li, Ruipeng Jia, Fang Fang 0009, Yanmin Shang, Yuhai Lu |
EMNLP (1) | 5 |
| 2021 | Multi-Granularity Heterogeneous Graph for Document-Level Relation ExtractionabstractReading text to extract relational facts has been a long-standing goal in natural language processing. It becomes especially challenging when the extraction scope is extended to document level, where multiple entities in a document generally exhibit complex intra- and inter-sentence relations. In this paper, we propose a novel Multi-granularity Heterogeneous Graph (MHG) to tackle this challenge. Specifically, we define four types of nodes with different granularities and eight types of edges based on heuristic rules, entrusting the MHG two major advantages. On the one hand, it connects any two entities with a short path in the graph to better handle the complex inter-sentence interactions between entities. On the other hand, it enables rich interactions among nodes with different granularities to promote accurate multi-hop reasoning. Experimental results on the largest document-level relation extraction dataset suggest that the proposed model achieves new state-of-the-art performance. Hengzhu Tang, Yanan Cao 0001, Zhenyu Zhang 0006, Ruipeng Jia, Fang Fang 0009, Shi Wang 0002 |
ICASSP | 5 |
| 2021 | No news is an island: Joint heterogeneous graph network for news classificationabstractNews classification is important for people to organize web information. There are many connections between daily news, for example, they may be different reports about the same event or the same person. However, previous works classify news mainly based on single news, they ignore relationships between multiple news. This paper innovatively utilizes relationships between multiple news, such as their relevance in time, place and people, to classify news. To take full advantage of these relationships and integrate various information of multiple news, we propose News Classification Graph (NCG), a heterogeneous graph with different types of nodes and edges. Furthermore, we propose Joint Heterogeneous graph Network (JHN) to properly embed the NCG. It fully utilizes the information of heterogeneous nodes and heterogeneous edges in NCG. Extensive experiments carried on four news datasets demonstrate the effectiveness of our work in solving the news classification problem. Zhezhou Kang, Yi Liu 0067, Guanqun Bi, Fang Fang 0009, Pengfei Yin |
ISCC | 4 |
| 2021 | Knowledge-Based Diverse Feature Transformation for Few-Shot Relation Classification
Yubao Tang, Zhezhou Li, Cong Cao 0001, Fang Fang 0009, Yanan Cao 0001, Yanbing Liu 0007, Jianhui Fu |
KSEM | 4 |
| 2020 | DistilSum: : Distilling the Knowledge for Extractive SummarizationabstractA popular choice for extractive summarization is to conceptualize it as sentence-level classification, supervised by binary labels. While the common metric ROUGE prefers to measure the text similarity, instead of the performance of classifier. For example, BERTSUMEXT, the best extractive classifier so far, only achieves a precision of 32.9% at the top 3 extracted sentences ([email protected]) on CNN/DM dataset. It is obvious that current approaches cannot model the complex relationship of sentences exactly with 0/1 targets. In this paper, we introduce DistilSum, which contains teacher mechanism and student model. Teacher mechanism produces high entropy soft targets at a high temperature. Our student model is trained with the same temperature to match these informative soft targets and tested with temperature of 1 to distill for ground-truth labels. Compared with large version of BERTSUMEXT, our experimental result on CNN/DM achieves a substantial improvement of 0.99 ROUGE-L score (text similarity) and 3.95 [email protected] score (performance of classifier). Our source code will be available on Github. Ruipeng Jia, Yanan Cao 0001, Haichao Shi, Fang Fang 0009, Yanbing Liu 0007, Jianlong Tan |
CIKM | 4 |
| 2020 | Neural Extractive Summarization with Hierarchical Attentive Heterogeneous Graph NetworkabstractSentence-level extractive text summarization is substantially a node classification task of network mining, adhering to the informative components and concise representations.There are lots of redundant phrases between extracted sentences, but it is difficult to model them exactly by the general supervised methods.Previous sentence encoders, especially BERT, specialize in modeling the relationship between source sentences.While, they have no ability to consider the overlaps of the target selected summary, and there are inherent dependencies among target labels of sentences.In this paper, we propose HAHSum (as shorthand for Hierarchical Attentive Heterogeneous Graph for Text Summarization), which well models different levels of information, including words and sentences, and spotlights redundancy dependencies between sentences.Our approach iteratively refines the sentence representations with redundancy-aware graph and delivers the label dependencies by message passing.Experiments on large scale benchmark corpus (CNN/DM, NYT, and NEWSROOM) demonstrate that HAHSum yields ground-breaking performance and outperforms previous extractive summarizers. Ruipeng Jia, Yanan Cao 0001, Hengzhu Tang, Fang Fang 0009, Cong Cao 0001, Shi Wang 0002 |
EMNLP (1) | 4 |
| 2020 | Enhancing Textual Representation for Abstractive Summarization: Leveraging Masked DecoderabstractFor existing models of abstractive summarization, the paradigm of autoregressive decoder inherently prefers relying on former tokens and the prediction error will propagate subsequently. To effectively eliminate the errors, we need a way to remodeling dependency during text generation. In this paper, we introduce MDSumma (as shorthand for Masked Decoder for Summarization), which masks partial tokens in decoder, aiming to alleviate the over-reliance on the antecedent. Moreover, with further facilitating the flexibility and diversity of textual representation, we employ a variational autoencoder model, sampling continuous latent variables from the probability distribution to explicitly model underlying semantics of the target summaries. Our architecture gives good balance between encoder contextual representation and decoder prediction, sidestepping the gap between training and inference. Experimental results on three benchmark datasets validate the effectiveness that our proposed method significantly outperforms the existing state-of-the-art approaches both on ROUGE and diversity scores. Ruipeng Jia, Yannan Cao, Fang Fang 0009, Jinpeng Li 0003, Yanbing Liu 0007, Pengfei Yin |
IJCNN | 3 |
| 2020 | Enhancing Pre-trained Language Representation for Multi-Task Learning of Scientific SummarizationabstractThis paper aims to extract summarization and keywords from scientific articles simultaneously, while abstract extraction (AE) and key extraction (KE) are considered as auxiliary tasks to each other. For the data scarcity in scientific AE and KE tasks, we propose a multi-task learning framework which uses huge unlabeled data to learn scientific language representation (pre-training) and uses smaller annotated data to transfer the learned representation to AE and KE (fine-tuning). Although the pre-trained language model performs well in universal natural language tasks, its capacity still has a margin of improvement for specific tasks. Inspired by this intuition, we use another two tasks keyword masking and key sentence prediction before the fine-tuning phase to enhance the language representation for AE and KE. This language representation enhancing stage uses the same labeled data but different optimization objectives with the fine-tuning phase. In order to evaluate our model, we develop and release a high-quality annotated corpus for scientific papers with keywords and abstract. We conduct comparative experiments on this dataset, and experimental results show that our multi-task learning framework achieves the state-of-the-art performance, proving the effectiveness of the language model enhancing mechanism. Ruipeng Jia, Yannan Cao, Fang Fang 0009, Jinpeng Li 0003, Yanbing Liu 0007, Pengfei Yin |
IJCNN | 3 |
| 2020 | HIN: Hierarchical Inference Network for Document-Level Relation Extraction
Hengzhu Tang, Yanan Cao 0001, Zhenyu Zhang 0006, Jiangxia Cao, Fang Fang 0009, Shi Wang 0002, Pengfei Yin |
PAKDD (1) | 5 |
| 2016 | Knowledge Extraction from Chinese Records of Cyber Attacks Based on a Semantic Grammar
Fang Fang 0009, Luchen Zhang, Cun-gen Cao 0001 |
KSEM | 1 |
| 2015 | A Chinese Framework of Semantic Taxonomy and Description: Preliminary Experimental Evaluation Using Web Information ExtractionabstractThe Chinese Framework of Semantic Taxonomy and Description (FSTD) is a linguistic resource that stores lexical and predicate-argument semantics about events or states in Chinese text, developed with the application of knowledge acquisition from Chinese text in mind. In this paper we build a web information extraction system, called NkiExtractor, to evaluate FSTD experimentally. We use two metrics: grammar coverage measures whether there is a semantic category of FSTD that corresponds to an event description in text, and extraction precision measures whether the correct predicate-argument structure can be extracted from text. Experimental results show that FSTD is a fairly comprehensive and effective resource for knowledge acquisition. We also discuss future work for expanding FSTD and improving extraction precision of NkiExtractor. Liangjun Zang, Weimin Wang 0002, Fang Fang 0009, Cong Cao 0001, Cun-gen Cao 0001 |
KSEM | 4 |