Fu Zhang 0001

dblp:03/4774-1 · DBLP profile ↗
← Back
61ranked-venue papers
29as first author
40since 2021 · last 2026
0000-0002-3880-8086ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 54 · 28 first-author · 34 since 2021Databases, data management, data science and information retrieval · 14 · 9 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 VCGD: Visual Clue Guided Decoding with Caption Model for Mitigating Hallucination in Multimodal Large Language Models
abstract
Multimodal large language models (MLLMs) demonstrate strong capabilities in multimodal understanding, reasoning, and interaction but still face the fundamental limitation of hallucinations, where they generate erroneous or fabricated information. Most existing research induces hallucinations by manually perturbing visual or instruction inputs, then uses output differences or model-generated descriptions as references to mitigate hallucinations and improve responsevisual consistency. However, these methods are constrained by model capabilities and prone to hallucination propagation. We propose Visual Clue Guided Decoding (VCGD), a novel decoding strategy that introduces an auxiliary Caption Model to generate precise visual clues during decoding for guiding model generation. It further incorporates image confidence constraints to critically suppress hallucination propagation during generation, thereby significantly improving content reliability and visual consistency. Specifically, VCGD leverages high-quality visual descriptions to guide MLLMs in correcting perceptual biases while generating answers. Furthermore, we introduce a Reinforcement Learning-based training paradigm for the Caption Model, in which a Reward Agent provides feedback on the quality of visual clues, further enhancing the accuracy of visual information. Extensive experiments across multiple benchmark datasets and state-of-the-art MLLMs demonstrate that VCGD significantly reduces hallucination rates and improves cross-modal consistency. Our method exhibits strong generalizability and scalability, offering an effective decoding enhancement strategy that can be seamlessly integrated into existing multimodal frameworks.
Fu Zhang 0001, Chenglong Lu, Jingwei Cheng
AAAI2
2026 A Boundary Token Graph for Zero-Shot Relation Triplet Extraction Involving Discontinuous Entities
abstract
Zero-Shot Relation Triplet Extraction (ZSRTE) aims to extract head-tail entity pairs and their corresponding relations from sentences, where the relations available during inference are not seen during training. Existing methods typically assume that entities are continuous; however, in practice, entities can be discontinuous, which poses challenges to these approaches. To address this issue, we are the first to discuss and study the ZSRTE task involving discontinuous entities, and propose an innovative BoG framework, which is based on our proposed Boundary Token Graph structure. This method first predicts and adds edges between boundary tokens of (dis)continuous entities to construct a token graph, and then innovatively transforms the relation triplet extraction task into a process of finding paths in the graph. Additionally, we design a Boundary Token-Aware Prompt for each relation to further enhance the interaction between boundary tokens and relation semantics. Experimental results on four ZSRTE datasets—with or without discontinuous entities—consistently demonstrate that our method outperforms previous approaches, achieving state-of-the-art results.
Kailun Lyu, Zehan Li, Fu Zhang 0001, Jingwei Cheng
AAAI3
2026 CamoQuery: Language-Guided Reasoning Camouflaged Object Segmentation
abstract
Although camouflaged object segmentation has advanced rapidly in recent years, existing methods are still confined to visual mask prediction under fixed task assumptions. They cannot interactively respond to user requests, nor can they proactively understand and reason about the user’s intent. Our work tackles this issue by proposing a novel task, Language-Guided Reasoning Camouflaged Object Segmentation (LRCOS). Given a camouflaged image and an implicit query text instruction that requires reasoning, LRCOS aims to output intent-consistent segmentation mask. To establish a benchmark for this task, we build CamoQuery, comprising 12,437 image–mask samples and 25971 implicit query text instructions. To better reflect real-world camouflaged scenarios, we additionally collect MCD, a multi-instance camouflage dataset where multiple camouflaged targets co-exist within the same scene, increasing the need for reasoning. Building on CamoQuery, we further propose COSA, a vision–language segmentation assistant that segments the intended camouflaged object from implicit queries and produces a reasoning explanation. Experiments on CamoQuery demonstrate that COSA has strong reasoning segmentation capability in camouflaged scenes and exhibits zero-shot capability.
Tianxin Han, Qing Dong 0004, Xingwei Wang 0001, Jie Jia 0001, Gang Wu 0007, Fu Zhang 0001
ACL (1)7
2026 Zero-shot Jianzi Recognition as Structured Visual Information Extraction in Open Compositional Symbolic Systems
abstract
Guqin (古 琴) Jianzi (減 字) is an open and freely compositional tablature system that encodes performance actions rather than acoustic outcomes.Its automatic recognition remains largely unexplored, as conventional OCR assumes a closed and enumerable glyph set and struggles with Jianzi's unbounded composition and manuscript-level variability.We introduce Zero-shot Jianzi Recognition, which formulates Jianzi recognition as visionto-sequence prediction of canonical component sequences under a zero-shot split.To enable scalable supervision, we construct Synthetic-JZ from aligned online composition metadata.We then synthesize manuscriptlike training images via component-wise style recomposition and manuscript-domain noise modeling, and fine-tune a VLM for end-toend component sequence recognition.At inference time, a lightweight legality-guided correction module re-ranks decoding candidates, suppressing structural hallucinations without modifying the backbone.Experiments on two benchmarks show that our method achieves 63.02% sequence accuracy on Real-JZ, our manually annotated realworld Jianzi benchmark, surpassing Gemini-3-Pro by 35.11%.This result highlights the feasibility of reliable automated Jianzi recognition and its potential for large-scale digitization of historical Guqin Jianzi Pu manuscripts.
Zehan Li, Fu Zhang 0001, Jingwei Cheng
ACL (1)2
2026 ATGL: An Adaptive-Threshold Global Loss for Document-level Relation Extraction
abstract
Document-level relation extraction (DocRE)aims to determine which relations hold between a given entity pair within a document.As a multi-label classification task, the most commonly adopted paradigm introduces a learnable threshold to distinguish positive and negative classes for an entity pair.Under this paradigm, existing losses decouple the optimization into independent positive and negative losses, which interact solely with a shared threshold.This leads to two inherent limitations: (i) threshold instability caused by conflicting gradient updates from the decoupled losses; and (ii) optimization bias exacerbated by the severe imbalance between limited positive samples and abundant negative samples inherent in DocRE, which makes the model more likely to predict that no relation exists.To address these issues, we propose the Adaptive-Threshold Global Loss (ATGL).Unlike prior work, ATGL integrates positive, negative, and threshold optimization into a unified logit space and explicitly enforces ranking constraints on their contributions to the objective.Furthermore, ATGL incorporates an imbalance-aware optimization mechanism, thereby effectively addressing the severe class imbalance in DocRE.Our ATGL serves as a general optimization objective that can be readily applied to different DocRE models.Experiments on four datasets show that ATGL outperforms other DocRE losses and achieves state-of-the-art results, while consistently improving the performance of existing DocRE models.
Huangming Xu, Fu Zhang 0001, Zhixuan Yang, Jingwei Cheng
ACL (1)2
2026 DEBAR: Mitigating Contextual Bias in Cross-Document Relation Extraction via Dual-Stream Decoupling
abstract
Cross-document Relation Extraction (CodRE) requires reasoning over scattered evidence to identify relations between target entities across multiple documents. Existing methods indiscriminately fuse target entities and the intermediate bridge entities that link them into a unified representation. This leads to intermediate evidence that often aligns with only one side of the entity pair, resulting in one-sided relation transfer contextual bias and incomplete reasoning chains. Moreover, these methods typically employ a global threshold to determine relation existence for all entity pairs, limiting the model’s reasoning performance.To address these issues, we propose DEBAR (Dual-stream Entity Bias Reduction), a framework designed to explicitly decouple and preserve bidirectional bridge evidence, combined with a novel dynamic loss optimization objective. Specifically, DEBAR employs a bridge-aware input construction strategy and a dual-stream graph reasoning network to separately encode head and tail contexts, preventing semantic interference while capturing global dependencies through iterative message passing. Furthermore, we introduce a curriculum-aware ranking optimization objective that progressively tightens classification constraints to stabilize training and enforce discriminative decision boundaries. Experiments on the CodRE benchmarks show that DEBAR achieves state-of-the-art performance while effectively mitigating cross-document contextual bias. Moreover, extensive experiments on our proposed loss across backbones confirm its generalization, suggesting it as a reliable replacement for existing CodRE losses. Code is available at https://github.com/newyuyou/DEBAR.
Zhixuan Yang, Fu Zhang 0001, Huangming Xu, Jingwei Cheng
ACL (1)2
2026 Countering Interest Over-Smoothing: Distilling Latent Factors via Diffusion for Multi-Interest Retrieval
abstract
Multi-interest recommendation is essential for the matching stage. By generating multiple user representations, it can better cover the diverse interests derived from user interaction history. Ideally, multi-interest models should effectively identify the underlying latent factors — the specific themes, intents, or preferences — within historical behaviors. However, conventional methods typically rely on weighted aggregation (e.g., Attention), which we argue leads to over-smoothed representations. This aggregation dilutes the intensity of significant patterns that appear only locally, blending them into a blurry average. To address this, we propose DMI, a model-agnostic diffusion framework that distills precise interests by amplifying co-occurring latent factors across behaviors. Distinct from prior diffusion works that reconstruct the single next item—which risks collapsing diverse interests—DMI reconstructs the interest vectors themselves to preserve their distributional independence. To support this, we introduce a cross-transformer module that adaptively extracts interest-specific information from designated historical interactions, transforming the diffusion process from an unconditional one into a guided, interest-disentangled pathway. In addition, we design a gradient back-propagation strategy to decouple the joint optimization of the reconstruction and recommendation losses, thereby improving training stability. Extensive offline experiments demonstrate DMI's superiority over existing methods, achieving an average relative improvement of 11.2% across all metrics on Amazon Books datasets while increasing recommendation diversity by 11.8%. Successfully deployed in a real-world recommender system, DMI effectively enhances user satisfaction and system performance at scale, serving the major traffic of hundreds of millions of daily active users.
Yankun Le, Fu Zhang 0001, Haoran Li 0011, Baoyuan Ou, Yingjie Qin, Zhixuan Yang, Ruilong Su
SIGIR2
2026 Dual-modal consistency learning for weakly supervised RGB-D camouflaged object detection with scribble annotations
Tianxin Han, Xingwei Wang 0001, Qing Dong 0004, Min Huang 0001, Jie Jia 0001, Fu Zhang 0001
Inf. Sci.7
2026 Temporal householder transformation embedding for temporal knowledge graph completion
Pengpeng Qiu, Fu Zhang 0001
Knowl. Based Syst.5
2026 Dual reasoning enhanced document-level relation extraction
Fu Zhang 0001, Yongxue Wu, Huangming Xu, Jingwei Cheng
Neural Comput. Appl.1
2025 Probing Relative Interaction and Dynamic Calibration in Multi-modal Entity Alignment
abstract
Multi-modal entity alignment aims to identify equivalent entities between two different multi-modal knowledge graphs.Current methods have made significant progress by improving embedding and cross-modal fusion.However, most of them depend on using loss functions to capture the relationship between modalities or adopt a one-time strategy to directly compute modality weights using attention mechanisms, which overlooks the relative interactions between modalities at the entity level and the accuracy of modality weights, thereby hindering the generalization to diverse entities.To address this challenge, we propose RICEA, a relative interaction and calibration framework for multi-modal entity alignment, which dynamically computes weights based on the relative interaction and recalibrates the weights according to their uncertainties.Among these, we propose a novel method called ADC that utilizes attention mechanisms to perceive the uncertainty of the weight for each modality, rather than directly calculating the weight of each modality as in previous works.Across 5 datasets and 23 settings, our proposed framework significantly outperforms other baselines.Our code and data are available at https://github.com/ChenxiaoLi-Joe/RICEA.
Chenxiao Li, Jingwei Cheng, Qiang Tong 0003, Fu Zhang 0001, Cairui Wang
ACL (1)4
2025 RRHF-V: Ranking Responses to Mitigate Hallucinations in Multimodal Large Language Models with Human Feedback
abstract
Multimodal large language models (MLLMs) demonstrate strong capabilities in multimodal understanding, reasoning, and interaction but still face the fundamental limitation of hallucinations, where they generate erroneous or fabricated information. To mitigate hallucinations, existing methods annotate pair-responses (one non-hallucination vs one hallucination) using manual methods or GPT-4V, and train alignment algorithms to improve the correspondence between images and text. More critically, an image description often involve multiple dimensions (e.g., object attributes, posture, and spatial relationships), making it challenging for the model to comprehensively learn multidimensional information from pair-responses. To this end, in this paper, we propose RRHFV, which is the first using rank-responses (one non-hallucination vs multiple ranking hallucinations) to mitigate multimodal hallucinations. Instead of using pair-responses to train the model, RRHF-V expands the number of hallucinatory responses, so that the responses with different scores in a rank-response enable the model to learn rich semantic information across various dimensions of the image. Further, we propose a scene graph-based approach to automatically construct rank-responses in a cost-effective and automatic manner. We also design a novel training objective based on rank loss and margin loss to balance the differences between hallucinatory responses within a rankresponse, thereby improving the model’s image comprehension. Experiments on two MLLMs of different sizes and four widely used benchmarks demonstrate that RRHF-V is effective in mitigating hallucinations and outperforms the DPO method based on pair-responses.
Fu Zhang 0001, Jinghao Lin, Chenglong Lu, Jingwei Cheng
COLING2
2025 SGMEA: Structure-Guided Multimodal Entity Alignment
abstract
Multimodal Entity Alignment (MMEA) aims to identify equivalent entities across different multimodal knowledge graphs (MMKGs) by integrating structural information, entity attributes, and visual data, thereby promoting knowledge sharing and deep multimodal data integration. However, existing methods often overlook the deeper connections between multimodal data. They primarily focus on the interactions between neighboring entities in the structural modality while neglecting the interactions between entities in the visual and attribute modalities. To address this, we propose a structure-guided multimodal entity alignment method (SGMEA), which prioritizes structural information from knowledge graphs to enhance the visual and attribute modalities. By fusing multimodal representations, SGMEA improves the accuracy of entity alignment. Experimental results demonstrate that SGMEA achieves stateof-the-art performance across multiple datasets, validating its effectiveness and superiority in practical applications.
Jingwei Cheng, Mingxiao Guo, Fu Zhang 0001
COLING3
2025 Exploring the Impacts of Feature Fusion Strategy in Multi-modal Entity Alignment
abstract
Multi-modal entity alignment aims to identify equivalent entities between two different multi-modal knowledge graphs, which consist of structural triples and images associated with entities. Unfortunately, prior works fuse the multi-modal knowledge of all entities only via solely one single fusion strategy. Therefore, the impact of the fusion strategy on individual entities could be largely ignored. To solve this challenge, we propose AMF2SEA, an adaptive multi-modal feature fusion strategy for entity alignment, which dynamically selects the optimal entity-level feature fusion strategy. Additionally, we build a new dataset based on DBP15K, which includes a full set of entity images from multiple inconsistent web sources, making it more representative of the real world. Experimental results demonstrate that our model achieves state-of-the-art (SOTA) performance compared to models using the same modality on DBP15K and its variants with richer image sources and styles. Our code and data are available at https://github.com/ChenxiaoLiJoe/AMFFSEA.
Chenxiao Li, Jingwei Cheng, Qiang Tong 0003, Fu Zhang 0001
COLING4
2025 Re-Cent: A Relation-Centric Framework for Joint Zero-Shot Relation Triplet Extraction
abstract
Zero-shot Relation Triplet Extraction (ZSRTE) aims to extract triplets from the context where the relation patterns are unseen during training. Due to the inherent challenges of the ZSRTE task, existing extractive ZSRTE methods often decompose it into named entity recognition and relation classification, which overlooks the interdependence of two tasks and may introduce error propagation. Motivated by the intuition that crucial entity attributes might be implicit in the relation labels, we propose a Relation-Centric joint ZSRTE method named Re-Cent. This approach uses minimal information, specifically unseen relation labels, to extract triplets in one go through a unified model. We develop two span-based extractors to identify the subjects and objects corresponding to relation labels, forming span-pairs. Additionally, we introduce a relation-based correction mechanism that further refines the triplets by calculating the relevance between span-pairs and relation labels. Experiments demonstrate that Re-Cent achieves state-of-the-art performance with fewer parameters and does not rely on synthetic data or manual labor.
Zehan Li, Fu Zhang 0001, Kailun Lyu, Jingwei Cheng, Tianyue Peng
COLING2
2025 DAEA: Enhancing Entity Alignment in Real-World Knowledge Graphs Through Multi-Source Domain Adaptation
abstract
Entity Alignment (EA) is a critical task in Knowledge Graph (KG) integration, aimed at identifying and matching equivalent entities that represent the same real-world objects. While EA methods based on knowledge representation learning have shown strong performance on synthetic benchmark datasets such as DBP15K, their effectiveness significantly decline in real-world scenarios which often involve data that is highly heterogeneous, incomplete, and domain-specific, as seen in datasets like DOREMUS and AGROLD. Addressing this challenge, we propose DAEA, a novel EA approach with Domain Adaptation that leverages the data characteristics of synthetic benchmarks for improved performance in real-world datasets. DAEA introduces a multi-source KGs selection mechanism and a specialized domain adaptive entity alignment loss function to bridge the gap between real-world data and optimal benchmark data, mitigating the challenges posed by aligning entities across highly heterogeneous KGs. Experimental results demonstrate that DAEA outperforms state-of-the-art models on real-world datasets, achieving a 29.94% improvement in Hits@1 on DOREMUS and a 5.64% improvement on AGROLD. Code is available at https://github.com/yangxiaoxiaoly/DAEA.
Linyan Yang, Shiqiao Zhou, Jingwei Cheng, Fu Zhang 0001, Jizheng Wan
COLING4
2025 CE-DA: Custom Embedding and Dynamic Aggregation for Zero-Shot Relation Extraction
abstract
Zero-shot Relation Extraction (ZSRE) aims to predict novel relations from sentences with given entity pairs, where the relations have not been encountered during training. Prototypebased methods, which achieve ZSRE by aligning the sentence representation and the relation prototype representation, have shown great potential. However, most existing works focus solely on improving the quality of prototype representations, neglecting sentence representations and lacking interaction between different types of relation side information. In this paper, we propose a novel ZSRE framework named CE-DA, which includes two modules: Custom Embedding and Dynamic Aggregation. We employ a two-stage approach to obtain customized embeddings of sentences. In the first stage, we train a sentence encoder through unsupervised contrastive learning, and in the second stage, we highlight the potential relations between entities in sentences using carefully designed entity emphasis prompts to further enhance sentence representations. Additionally, our dynamic aggregation method assigns different weights to different types of relation side information through a learnable network to enhance the quality of relation prototype representations. In contrast to traditional methods that treat the importance of all side information equally, our dynamic aggregation method further strengthen the interaction between different types of relation side information. Our method demonstrates competitive performance across various metrics on two ZSRE datasets.
Fu Zhang 0001, Zehan Li, Jingwei Cheng
COLING1
2025 A Dual-Task Learning Model for Temporal Knowledge Graph Entity Alignment
Jingwei Cheng, Xihao Wang, Fu Zhang 0001
DASFAA (3)3
2025 Frame First, Then Extract: A Frame-Semantic Reasoning Pipeline for Zero-Shot Relation Triplet Extraction
abstract
Large Language Models (LLMs) have shown impressive capabilities in language understanding and generation, leading to growing interest in zero-shot relation triplet extraction (Ze-roRTE), a task that aims to extract triplets for unseen relations without annotated data.However, existing methods typically depend on costly fine-tuning and lack the structured semantic guidance required for accurate and interpretable extraction.To overcome these limitations, we propose FrameRTE, a novel Ze-roRTE framework that adopts a "frame first, then extract" paradigm.Rather than extracting triplets directly, FrameRTE first constructs high-quality Relation Semantic Frames (RSFs) through a unified pipeline that integrates frame retrieval, synthesis, and enhancement.These RSFs serve as structured and interpretable knowledge scaffolds that guide frozen LLMs in the extraction process.Building upon these RSFs, we further introduce a human-inspired three-stage reasoning pipeline consisting of semantic frame evocation, frame-guided triplet extraction, and core frame elements validation to achieve semantically constrained extraction.Experiments demonstrate that FrameRTE achieves competitive zero-shot performance on multiple benchmarks.Moreover, the RSFs we construct serve as high-quality semantic resources that can enhance other extraction methods, showcasing the synergy between linguistic knowledge and foundation models. Frame ( Work )An Agent expends effort towards achieving a Goal.Alternatively, a Salient_entity involved in the Goal can be expressed in place of a Goal expression.Definition Agent: The Agent puts effort into reaching Goal. Core Frame Elements Goal:The Goal is what the Agent expends effort to achieve.Salient_entity: An entity that is centrally involved in the Goal that the Agent is attempting to acheive.Circumstances, Degree,
Zehan Li, Fu Zhang 0001, Jingwei Cheng, Tianyue Peng
EMNLP2
2025 Multi-Frequency Contrastive Decoding: Alleviating Hallucinations for Large Vision-Language Models
abstract
Large visual-language models (LVLMs) have demonstrated remarkable performance in visual-language tasks.However, object hallucination remains a significant challenge for LVLMs.Existing studies attribute object hallucinations in LVLMs mainly to linguistic priors and data biases.We further explore the causes of object hallucinations from the perspective of frequency domain and reveal that insufficient frequency information in images amplifies these linguistic priors, increasing the likelihood of hallucinations.To mitigate this issue, we propose the Multi-Frequency Contrastive Decoding (MFCD) method, a simple yet trainingfree approach that removes the hallucination distribution in the original output distribution, which arises from LVLMs neglecting the highfrequency information or low-frequency information in the image input.Without compromising the general capabilities of LVLMs, the proposed MFCD effectively mitigates the object hallucinations in LVLMs.Our experiments demonstrate that MFCD significantly mitigates object hallucination across diverse large-scale vision-language models, without requiring additional training or external tools.In addition, MFCD can be applied to various LVLMs without modifying model architecture or requiring additional training, demonstrating its generality and robustness.Codes are available at https://github.com/liubq-dev/mfcd.
Fu Zhang 0001, Jingwei Cheng
EMNLP2
2025 Breaking the Noise Barrier: LLM-Guided Semantic Filtering and Enhancement for Multi-Modal Entity Alignment
abstract
Multi-modal entity alignment (MMEA) aims to identify equivalent entities between two multimodal knowledge graphs (MMKGs).Existing methods have made substantial advancements in enhancing multi-modal fusion.However, the intrinsic noise within modalities, such as the inconsistency in visual modality and redundant attributes, has not been thoroughly investigated.Excessive noise not only weakens semantic representation but also increases the risk of overfitting in attention-based fusion methods.To address this, we propose LGEA (LLM-Guided Entity Alignment), a novel LLM-guided MMEA framework that prioritizes noise reduction before fusion.Specifically, LGEA introduces two key strategies: (1) fine-grained visual filtering to remove irrelevant images at the semantic level, and (2) contextual summarization of attribute information to enhance entity semantics.To our knowledge, we are the first work to apply LLMs for both visual filtering and attribute-level semantic enhancement in MMEA.Experiments on multiple benchmarks, including the noisy FBYG dataset, show that LGEA sets a new state-of-the-art (SOTA) in robust multi-modal alignment, highlighting the potential of noiseaware strategies as a promising direction for future MMEA research 1 .
Chenglong Lu, Chenxiao Li, Jingwei Cheng, Yongquan Ji, Fu Zhang 0001
EMNLP6
2025 ARPDL: Adaptive Relational Prior Distribution Loss as an Adapter for Document-Level Relation Extraction
abstract
The goal of document-level relation extraction (DocRE) is to identify relations between entities from multiple sentences. As a multi-label classification task, a common approach is to determine whether there are relations for an entity pair by selecting a multi-label classification threshold, with scores of relations above the threshold predicted as positive and the rest as negative. However, we find that predicting multiple relations for entity pairs causes the decrease of predicted scores in positive classes. This could lead to many positive classes being incorrectly predicted as negative. Additionally, our analysis suggests that fitting the distribution of predicted relations to the prior distribution of relations can help improve prediction performance. However, previous studies have not explored or leveraged the prior distribution of relations. To address these issues and findings, we for the first time propose the idea of incorporating the relational prior distribution into the loss calculation in DocRE tasks. We innovatively propose an Adaptive Relational Prior Distribution Loss (ARPDL), which can adaptively adjust relation prediction scores based on the relational prior distribution. Our designed relational prior distribution component can also be integrated as an adapter into other threshold-based losses to improve prediction performance. Experimental results demonstrate that ARPDL consistently improves the performance of existing DocRE models, achieving new state-of-the-art results. Furthermore, integrating our relational prior distribution adapter into other losses significantly enhances their performance in DocRE tasks, validating the effectiveness and generality of our approach. Code is available at https://github.com/xhm-code/ARPDL.
Huangming Xu, Fu Zhang 0001, Jingwei Cheng
IJCAI2
2025 Rethinking the Role of LLMs for Document-level Relation Extraction: a Refiner with Task Distribution and Probability Fusion
abstract
Fu Zhang, Xinlong Jin, Jingwei Cheng, Hongsen Yu, Huangming Xu. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Fu Zhang 0001, Xinlong Jin, Jingwei Cheng, Hongsen Yu, Huangming Xu
NAACL (Long Papers)1
2025 Industrial device-aided data collection for real-time rail defect detection via a lightweight network
Qing Dong 0004, Tianxin Han, Gang Wu 0007, Min Huang 0001, Fu Zhang 0001
Eng. Appl. Artif. Intell.6
2025 A self-supervised method for learning path-augmented knowledge graph embedding
Fu Zhang 0001, Jingwei Cheng
Eng. Appl. Artif. Intell.2
2025 Mention Distance-aware Interactive Attention with Multi-step Reasoning for document-level relation extraction
Fu Zhang 0001, Huangming Xu, Jingwei Cheng
Eng. Appl. Artif. Intell.1
2025 Weakly supervised camouflaged object detection as Progressive Perception Learning
Tianxin Han, Xingwei Wang 0001, Qing Dong 0004, Min Huang 0001, Jie Jia 0001, Fu Zhang 0001
Knowl. Based Syst.6
2024 SRF: Enhancing Document-Level Relation Extraction with a Novel Secondary Reasoning Framework
abstract
Document-level Relation Extraction (DocRE) aims to extract relations between entity pairs in a document and poses many challenges as it involves multiple mentions of entities and crosssentence inference.However, several aspects that are important for DocRE have not been considered and explored.Existing work ignore bidirectional mention interaction when generating relational features for entity pairs.Also, sophisticated neural networks are typically designed for cross-sentence evidence extraction to further enhance DocRE.More interestingly, we reveal a noteworthy finding: If a model has predicted a relation between an entity and other entities, this relation information may help infer and predict more relations between the entity's adjacent entities and these other entities.Nonetheless, none of existing methods leverage secondary reasoning to exploit results of relation prediction.To this end, we propose a novel Secondary Reasoning Framework (SRF) for DocRE.In SRF, we initially propose a DocRE model that incorporates bidirectional mention fusion and a simple yet effective evidence extraction module (incurring only an additional learnable parameter overhead) for relation prediction.Further, for the first time, we elaborately design and propose a novel secondary reasoning method to discover more relations by exploring the results of the first relation prediction.Extensive experiments show that SRF achieves SOTA performance and our secondary reasoning method is both effective and general when integrated into existing models. 1
Fu Zhang 0001, Qi Miao, Jingwei Cheng, Hongsen Yu, Yongxue Wu
EMNLP1
2024 ATAP: Automatic Template-Augmented Commonsense Knowledge Graph Completion via Pre-Trained Language Models
abstract
The mission of commonsense knowledge graph completion (CKGC) is to infer missing facts from known commonsense knowledge.CKGC methods can be roughly divided into two categories: triple-based methods and text-based methods.Due to the imbalanced distribution of entities and limited structural information, triple-based methods struggle with long-tail entities.Text-based methods alleviate this issue, but require extensive training and fine-tuning of language models, which reduces efficiency.To alleviate these problems, we propose ATAP, the first CKGC framework that utilizes automatically generated continuous prompt templates combined with pre-trained language models (PLMs).Moreover, ATAP uses a carefully designed new prompt template training strategy, guiding PLMs to generate optimal prompt templates for CKGC tasks.Combining the rich knowledge of PLMs with the template automatic augmentation strategy, ATAP effectively mitigates the long-tail problem and enhances CKGC performance.Results on benchmark datasets show that ATAP achieves state-of-theart performance overall. 1
Fu Zhang 0001, Jingwei Cheng
EMNLP1
2024 Attr-Int: A Simple and Effective Entity Alignment Framework for Heterogeneous Knowledge Graphs
abstract
Entity alignment (EA) refers to the task of linking entities in different knowledge graphs (KGs). Existing EA methods rely heavily on structural isomorphism. However, in real-world KGs, aligned entities usually have non-isomorphic neighborhood structures, which paralyses the application of these structure-dependent methods. In this paper, we investigate and tackle the problem of entity alignment between heterogeneous KGs. First, we propose two new benchmarks to closely simulate real-world EA scenarios of heterogeneity. Then we conduct extensive experiments to evaluate the performance of representative EA methods on the new benchmarks. Finally, we propose a simple and effective entity alignment framework called Attr-Int, in which innovative attribute information interaction methods can be seamlessly integrated with any embedding encoder for entity alignment, improving the performance of existing entity alignment techniques. Experiments demonstrate that our framework outperforms the state-of-the-art approaches on two new benchmarks.
Linyan Yang, Jingwei Cheng, Chuanhao Xu, Xihao Wang, Fu Zhang 0001
ICASSP6
2024 gMLP-KGE: a simple but efficient MLPs with gating architecture for link prediction
Fu Zhang 0001, Pengpeng Qiu, Jingwei Cheng
Appl. Intell.1
2024 A bitemporal RDF index based on skip list
abstract
The Resource Description Framework (RDF) is a framework for expressing information about resources in the form of triples (subject, predicate, object). The information represented by the standard RDF is static, i.e., that does not change over time. To better deal with a large amount of time-related information, temporal RDF is proposed. Consequently, how to explore index technology to efficiently query temporal information has become an important research issue, but the research on the index of temporal RDF is still short, especially the index of bitemporal RDF. BitemporalRDF can represent more complicated situations (e.g., RDF triples with both validtime and transactiontime). Indexes for bitemporal RDF can further expand the application scenarios and functions of temporal RDF. In this paper, we propose an efficient index for bitemporal RDF queries. The index innovatively introduces and re-designs skip list structure into the bitemporal RDF query. We also investigate how to cover almost all query patterns with as few indexes as possible. In addition, although the proposed index is conceived for temporal RDF, it also takes into account the performance of standard RDF queries when the time element is unknown. Finally, we run experiments with synthetic data sets of different sizes using the Lehigh University Benchmark (LUBM), and results prove that the proposed index is scalable and effective.
Fu Zhang 0001
Intell. Data Anal.1
2024 Joint framework for tensor decomposition-based temporal knowledge graph completion
Fu Zhang 0001, Yuzhe Shi, Jingwei Cheng, Jinghao Lin
Inf. Sci.1
2023 Multi-Aspect Enhanced Convolutional Neural Networks for Knowledge Graph Completion
abstract
Knowledge graph completion (KGC, also referred to as link prediction) aims at predicting missing entities and relations in knowledge graphs (KGs). Knowledge graph embedding (KGE) techniques have been proven to be effective for link prediction. Currently, a series of convolutional neural networks (CNNs) based models (e.g., ConvE and its extended models) have attained excellent results for link prediction. However, several aspects that are important for link prediction using CNNs have not been considered and enhanced simultaneously, which significantly limit the performance of these models. In this paper we explore an effective KGE model based on CNNs. We investigate and discover four extremely important aspects that have a strong influence on ConvE: entity and relation embeddings, entity-to-relation interaction approaches, CNN structure, and loss function. Based on the optimization of the above four aspects, we propose a novel KGE method called ConvEICF. Through extensive experiments, we find that ConvEICF outperforms the previous state-of-the-art link prediction baselines on FB15k-237 and WN18RR datasets. In particular, ConvEICF achieves a Hits@10 score that is 11.2% and 6.5% better than ConvE on FB15k-237 and WN18RR datasets respectively. Additionally, through in-depth experiments we observe an interesting phenomenon and important finding that the very common 1-N scoring technique in KGE can be considerably improved by just adding a dropout operation. Our code is available at https://github.com/NEU-IDKE/ConvEICF.
Fu Zhang 0001, Pengpeng Qiu, Jingwei Cheng
ECAI1
2023 Semi-Supervised Semantic Segmentation with Structured Output Space Adaption
abstract
Semi-supervised semantic segmentation methods rely on dense pixel-level classification with limited data and can thus be developed to adapt source ground truth labels to a target domain. In this paper, we creatively propose a method for semi-supervised semantic segmentation. The key innovation is our adversarial learning method for space adaptation in context, which can be regarded as a structured output that contains spatial similarities between unlabeled data and labeled data. To achieve this, we construct an adversarial learning network to efficiently adapt to the structural output space in labeled and unlabeled similar samples. Furthermore, we introduce two learning strategies into semi-supervised semantic segmentation, one that can selectively capture intra-category and inter-category context dependencies, resulting in robust feature representations. While the other explicitly concatenates the shape information of objects as a separate processing branch to produce sharper predicted boundaries of objects. Experimental results on two well-known benchmark datasets show that our method achieves better performance compared to the previous competitive models.
Weiquan Huang, Fu Zhang 0001
ICASSP2
2023 Improving entity alignment via attribute and external knowledge filtering
Fu Zhang 0001, Jingwei Cheng
Appl. Intell.1
2023 A joint training network for learning more distinguishable relation features in relation classification
Fu Zhang 0001, Jiejie Qin, Jingwei Cheng
Knowl. Based Syst.1
2022 Constructing ontologies by mining deep semantics from XML Schemas and XML instance documents
abstract
With the development of the Semantic Web and Artificial Intelligence techniques, ontology has become a very powerful way of representing not only knowledge but also their semantics. Therefore, how to construct ontologies from existing data sources has become an important research topic. In this paper, an approach for constructing ontologies by mining deep semantics from eXtensible Markup Language (XML) Schemas (including XML Schema 1.0 and XML Schema 1.1) and XML instance documents is proposed. Given an XML Schema and its corresponding XML instance document, 34 rules are first defined to mine deep semantics from the XML Schema. The mined semantics is formally stored in an intermediate conceptual model and then is used to generate an ontology at the conceptual level. Further, an ontology population approach at the instance level based on the XML instance document is proposed. Now, a complete ontology is formed. Also, some corresponding core algorithms are provided. Finally, a prototype system is implemented, which can automatically generate ontologies from XML Schemas and populate ontologies from XML instance documents. The paper also classifies and summarizes the existing work and makes a detailed comparison. Case studies on real XML data sets verify the effectiveness of the approach.
Fu Zhang 0001
Int. J. Intell. Syst.1
2022 An MRC and adaptive positive-unlabeled learning framework for incompletely labeled named entity recognition
abstract
Currently, named entity recognition (NER) is mainly evaluated on standard and well-annotated data sets. However, the construction of a well-annotated data set will consume a lot of manpower and time. In lots of applications of NER, data sets may contain a lot of noise, and a large part of noise comes from unlabeled entities. At present, the training process of most models treat unlabeled entities as nonentities, which causes these models to lean toward predicting most words of an input context as nonentities and greatly affects their performances. In this paper, as the first attempt, we innovatively propose an adaptive positive-unlabeled (adaPU) learning technology, and integrate the adaPU into a machine reading comprehension (MRC) framework for NER, which can still perform well on data sets with a large proportion of unlabeled entities. In our framework, to leverage the above problem that a model may predict most words of an input context as nonentities, we propose an adaPU learning technology by adjusting a loss coefficient of positive and negative samples. Moreover, instead of just constructing a fixed query for each entity type as input to MRC, we propose a new method of dynamically constructing multiple queries for each entity type, which also brings slight performance improvement for NER. Accordingly, we explore new training and entity inference strategies for our learning framework. The experimental results show that our framework is effective on data sets that contain a large number of unlabeled entities. When the proportion of unlabeled entities reaches 50%, our framework still can keep from losing effectiveness and maintain more than 80 F1-scores on several data sets. Also, the experiments show that our framework can achieve better or competitive performance on standard data sets. The ablation experiments further fully demonstrate our MRC framework with adaPU learning and dynamic query construction method can improve the performance of NER.
Fu Zhang 0001, Liangdong Ma, Jingwei Cheng
Int. J. Intell. Syst.1
2022 A comprehensive overview of knowledge graph completion
Fu Zhang 0001, Jingwei Cheng
Knowl. Based Syst.2
2019 Representation Learning of Knowledge Graphs with Multi-scale Capsule Network
Jingwei Cheng, Jinming Dang, Chunguang Pan, Fu Zhang 0001
IDEAL (1)5
2018 Storing fuzzy description logic ontology knowledge bases in fuzzy relational databases
Fu Zhang 0001, Zongmin Ma 0001, Qiang Tong 0003, Jingwei Cheng
Appl. Intell.1
2016 Enhanced entity-relationship modeling with description logic
Fu Zhang 0001, Zongmin Ma 0001, Jingwei Cheng
Knowl. Based Syst.1
2015 Storing OWL ontologies in object-oriented databases
Fu Zhang 0001, Zongmin Ma 0001
Knowl. Based Syst.1
2014 Representing and Reasoning About XML with Ontologies
Fu Zhang 0001, Zongmin Ma 0001
Appl. Intell.1
2013 Construction of fuzzy OWL ontologies from fuzzy EER models: A semantics-preserving approach
Fu Zhang 0001, Zongmin Ma 0001, Li Yan 0001, Jingwei Cheng
Fuzzy Sets Syst.1
2013 Construction of fuzzy ontologies from fuzzy XML models
Fu Zhang 0001, Zongmin Ma 0001, Li Yan 0001
Knowl. Based Syst.1
2012 A fuzzy ontology approach for representing Fuzzy Petri Nets
abstract
Petri Net (PN) has proven to be quite effective tool for graphical modeling, mathematical modeling, simulation, and real time control by the use of places and transitions. However, information imprecision and uncertainty exist in many real-world applications, and PNs found themselves inadequate to address the problems of imprecision and uncertainty in data. Therefore, the Fuzzy Petri Net (FPN) was developed and had been employed in many different fields like communication, manufacturing, electronics, and etc. In particular, with the wide utilization of FPNs, many researchers suggest that FPNs should be reused and shared. Emerging the Semantic Web technologies, such as fuzzy ontology, can play an important role in this scenario. On this basis, in this paper, we propose a fuzzy ontology approach for representing FPNs. First, we propose a formal definition of FPNs. Then, we give a complete definition of fuzzy OWL ontologies, where fuzzy ontologies formulated in fuzzy OWL language are called fuzzy OWL ontologies. Based on the formalization of FPNs and fuzzy OWL ontologies, we further propose a fuzzy ontology approach for representing FPNs, where we translate the key features of FPNs into the elements of fuzzy OWL ontologies such as fuzzy classes, fuzzy properties, fuzzy individuals, and fuzzy axioms. Finally, based on the translated fuzzy ontologies, we briefly discuss how to reason on FPNs through the reasoning mechanism of fuzzy ontologies.
Fu Zhang 0001, Zongmin Ma 0001, Li Yan 0001
FUZZ-IEEE1
2012 Modeling fuzzy information in UML class diagrams and object-oriented database models
Zongmin Ma 0001, Li Yan 0001, Fu Zhang 0001
Fuzzy Sets Syst.3
2012 A description logic approach for representing and reasoning on fuzzy object-oriented database models
Fu Zhang 0001, Zongmin Ma 0001, Li Yan 0001, Yu Wang 0054
Fuzzy Sets Syst.1
2012 Reasoning of fuzzy relational databases with fuzzy ontologies
abstract
A significant interest developed regarding the problem of describing databases with expressive knowledge representation techniques in recent years, so that database reasoning may be handled intelligently. Therefore, it is possible and meaningful to investigate how to reason on fuzzy relational databases (FRDBs) with fuzzy ontologies. In this paper, we first propose a formal approach and an automated tool for constructing fuzzy ontologies from FRDBs, and then we study how to reason on FRDBs with constructed fuzzy ontologies. First, we give their respective formal definitions of FRDBs and fuzzy Web Ontology Language (OWL) ontologies. On the basis of this, we propose a formal approach that can directly transform an FRDB (including its schema and data information) into a fuzzy OWL ontology (consisting of the fuzzy ontology structure and instance). Furthermore, following the proposed approach, we implement a prototype construction tool called FRDB2FOnto. Finally, based on the constructed fuzzy OWL ontologies, we investigate how to reason on FRDBs (e.g., consistency, satisfiability, subsumption, and redundancy) through the reasoning mechanism of fuzzy OWL ontologies, so that the reasoning of FRDBs may be done automatically by means of the existing fuzzy ontology reasoner.© 2012 Wiley Periodicals, Inc.
Fu Zhang 0001, Li Yan 0001, Zongmin Ma 0001
Int. J. Intell. Syst.1
2011 Storing Fuzzy Ontology in Fuzzy Relational Database
Fu Zhang 0001, Zongmin Ma 0001, Li Yan 0001, Jingwei Cheng
DEXA (2)1
2011 Representing and reasoning on fuzzy UML models: A description logic approach
Zongmin Ma 0001, Fu Zhang 0001, Li Yan 0001, Jingwei Cheng
Expert Syst. Appl.2
2010 Formal approach and automated tool for constructing ontology from object-oriented database model
abstract
Extracting domain knowledge from databases can facilitate the development of Web ontologies. In this paper, a formal approach and an automated tool for constructing ontologies from Object-oriented database models (OODMs) are developed. The approach and tool can automatically translate an OODM and its corresponding database instances into the ontology structure and ontology instances, respectively. Case studies show that the approach is feasible and the automated construction tool is efficient.
Fu Zhang 0001, Zongmin Ma 0001, Xing Wang 0002, Yu Wang 0054
CIKM1
2010 RIF Centered Rule Interchange in the Semantic Web
Xing Wang 0002, Zongmin Ma 0001, Fu Zhang 0001, Li Yan 0001
DEXA (1)3
2010 Automatic Fuzzy Semantic Web Ontology Learning from Fuzzy Object-Oriented Database Model
Fu Zhang 0001, Zongmin Ma 0001, Gaofeng Fan, Xing Wang 0002
DEXA (1)1
2010 Formal semantics-preserving translation from fuzzy ER model to fuzzy OWL DL ontology
abstract
Ontology is an important part of the W3C standards for the Semantic Web, and how to quickly and cheaply construct Web ontologies has become a key technology to enable the Semantic Web. However, information imprecision and uncertainty exist in many re
Zongmin Ma 0001, Fu Zhang 0001, Li Yan 0001, Yanhui Lv
Web Intell. Agent Syst.2
2009 Fuzzy semantic web ontology learning from fuzzy UML model
abstract
How to quickly and cheaply construct Web ontologies has become a key technology to enable the Semantic Web. Classical ontologies are not sufficient for handling imprecise and uncertain information that is commonly found in many application domains. In this paper, we propose an approach for constructing fuzzy ontologies from fuzzy UML models, in which the fuzzy ontology consists of fuzzy ontology structure and instances. Firstly, the fuzzy UML model is investigated in detail, and a kind of formal definition of fuzzy UML models is proposed. Then, a kind of fuzzy ontology called fuzzy OWL DL ontology is introduced. Furthermore, we consider the fuzzy UML model and the corresponding fuzzy UML instantiations (i.e., object diagrams) simultaneously, and translate them into the fuzzy ontology structure and the fuzzy ontology instances, respectively. In addition, since a fuzzy OWL DL ontology is equivalent to a fuzzy Description Logic f-SHOIN(D) knowledge base, how the reasoning problems of fuzzy UML models (e.g., consistency, subsumption, equivalence, and redundancy) may be reasoned through reasoning mechanism of f-SHOIN(D) is investigated, which can help to construct fuzzy ontologies more exactly.
Fu Zhang 0001, Zongmin Ma 0001, Jingwei Cheng, Xiangfu Meng
CIKM1
2009 Deciding Query Entailment in Fuzzy Description Logic Knowledge Bases
Jingwei Cheng, Zongmin Ma 0001, Fu Zhang 0001, Xing Wang 0002
DEXA3
2008 Representation and reasoning of fuzzy ER model with description logic
abstract
Information imprecision and uncertainty exist in many real-world applications and hence fuzzy data modeling has been extensively investigated in various data models. This paper focuses on the representation and reasoning of fuzzy ER data model with description logic. Firstly, we give the formal definition and semantics of fuzzy ER model. Then based on the description logic DLR, a kind of new fuzzy description logic, i.e., fuzzy description logic FDLR (fuzzy DLR), is presented thoroughly. The definitions of syntax, semantics, and knowledge base form are given for FDLR. The fuzzy ER model with fuzzy description logic FDLR is investigated to translate fuzzy ER model into FDLR knowledge bases. With an example, the fact that the fuzzy ER model can be well represented by FDLR can be explained. The reasoning problem of satisfiability, subsumption relation, and redundancy of fuzzy ER model may reason automatically through reasoning mechanism of fuzzy description logic FDLR. The correctness of translation and reasoning problems are proved.
Fu Zhang 0001, Zongmin Ma 0001, Li Yan 0001
FUZZ-IEEE1
2008 Formal Semantics-Preserving Translation from Fuzzy ER Model to Fuzzy OWL DL Ontology
abstract
How to quickly and cheaply construct Web ontologies has become a key technology to enable the Semantic Web. However, information imprecision and uncertainty exist in many real-world applications. Thus constructing fuzzy ontology by extracting domain knowledge from fuzzy database model such as fuzzy ER model can profitably support fuzzy ontology development. In this paper, firstly, we give the formal definition and semantics of fuzzy ER model. Then, we introduce a kind of fuzzy extension of OWL DL, named fuzzy OWL DL. Furthermore, based on the fuzzy OWL DL, the formal definition and Model-Theoretic semantics of fuzzy OWL DL ontology are given. Whatpsilas more, we realize the formal translation from fuzzy ER model to fuzzy OWL DL ontology by a semantics-preserving translation algorithm. Finally, since a fuzzy OWL DL ontology is being equivalent to a description logic f-SHOIN(D) knowledge base, the reasoning problem of satisfiability, subsumption, and redundancy of fuzzy ER model may reason automatically through reasoning mechanism of f-SHOIN(D) is also investigated, which can contribute to constructing fuzzy OWL DL ontologys exactly that meet applicationpsilas needs.
Fu Zhang 0001, Zongmin Ma 0001, Yanhui Lv, Xing Wang 0002
Web Intelligence1