Zhenyu Zhang 0006

dblp:01/1844-6 · DBLP profile ↗
← Back
37ranked-venue papers
7as first author
21since 2021 · last 2026
0000-0002-5936-6678ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 27 · 4 first-author · 17 since 2021Databases, data management, data science and information retrieval · 11 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Uncertainty-Aware Routing for Principled Alignment with MoE Dynamics
abstract
Yilong Chen, Junyuan Shang, Yuchen Feng, Zhenyu Zhang, Naibin Gu, Ziqi Wang, Tingwen Liu, Shuohuan Wang, Yu Sun, Hua Wu, Haifeng Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Junyuan Shang, Zhenyu Zhang 0006, Naibin Gu, Tingwen Liu, Shuohuan Wang, Hua Wu 0003, Haifeng Wang 0001
ACL (1)4
2025 Inner Thinking Transformer: Leveraging Dynamic Depth Scaling to Foster Adaptive Internal Thinking
abstract
Yilong Chen, Junyuan Shang, Zhenyu Zhang, Yanxi Xie, Jiawei Sheng, Tingwen Liu, Shuohuan Wang, Yu Sun, Hua Wu, Haifeng Wang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Junyuan Shang, Zhenyu Zhang 0006, Yanxi Xie, Jiawei Sheng, Tingwen Liu, Shuohuan Wang, Yu Sun 0029, Hua Wu 0003, Haifeng Wang 0001
ACL (1)3
2025 BeamLoRA: Beam-Constraint Low-Rank Adaptation
abstract
Due to the demand for efficient fine-tuning of large language models, Low-Rank Adaptation (LoRA) has been widely adopted as one of the most effective parameter-efficient fine-tuning methods. Nevertheless, while LoRA improves efficiency, there remains room for improvement in accuracy. Herein, we adopt a novel perspective to assess the characteristics of LoRA ranks. The results reveal that different ranks within the LoRA modules not only exhibit varying levels of importance but also evolve dynamically throughout the fine-tuning process, which may limit the performance of LoRA. Based on these findings, we propose BeamLoRA, which conceptualizes each LoRA module as a beam where each rank naturally corresponds to a potential sub-solution, and the fine-tuning process becomes a search for the optimal sub-solution combination. BeamLoRA dynamically eliminates underperforming sub-solutions while expanding the parameter space for promising ones, enhancing performance with a fixed rank. Extensive experiments across three base models and 12 datasets spanning math reasoning, code generation, and commonsense reasoning demonstrate that BeamLoRA consistently enhances the performance of LoRA, surpassing the other baseline methods.
Naibin Gu, Zhenyu Zhang 0006, Xiyu Liu 0003, Peng Fu 0008, Zheng Lin 0001, Shuohuan Wang, Hua Wu 0003, Weiping Wang 0005, Haifeng Wang 0001
ACL (1)2
2025 HFT: Half Fine-Tuning for Large Language Models
abstract
Large language models (LLMs) with one or more fine-tuning phases have become necessary to unlock various capabilities, enabling LLMs to follow natural language instructions and align with human preferences. However, it carries the risk of catastrophic forgetting during sequential training, the parametric knowledge or the ability learned in previous stages may be overwhelmed by incoming training data. This paper finds that LLMs can restore some original knowledge by regularly resetting partial parameters. Inspired by this, we introduce Half Fine-Tuning (HFT) for LLMs, as a substitute for full fine-tuning (FFT), to mitigate the forgetting issues, where half of the parameters are selected to learn new tasks. In contrast, the other half are frozen to retain previous knowledge. We provide a feasibility analysis from the optimization perspective and interpret the parameter selection operation as a regularization term. HFT could be seamlessly integrated into existing fine-tuning frameworks without changing the model architecture. Extensive experiments and analysis on supervised fine-tuning, direct preference optimization, and continual learning consistently demonstrate the effectiveness, robustness, and efficiency of HFT. Compared with FFT, HFT not only significantly alleviates the forgetting problem, but also achieves the best performance in a series of downstream benchmarks, with an approximately 30% reduction in training time.
Tingfeng Hui, Zhenyu Zhang 0006, Shuohuan Wang, Weiran Xu, Yu Sun 0029, Hua Wu 0003
ACL (1)2
2025 Upcycling Instruction Tuning from Dense to Mixture-of-Experts via Parameter Merging
abstract
Mixture-of-Experts (MoE) shines brightly in large language models (LLMs) and demonstrates outstanding performance in plentiful natural language processing tasks.However, existing methods transforming LLMs from dense to MoE face significant data requirements and typically rely on large-scale post-training.In this paper, we propose Upcycling Instruction Tuning (UpIT), a data-efficient approach for tuning a dense pre-trained model into a MoE instruction model.Specifically, we first point out that intermediate checkpoints during instruction tuning of the dense model are naturally suitable for specialized experts, and then propose an expert expansion stage to flexibly achieve models with flexible numbers of experts, where genetic algorithm and parameter merging are introduced to ensure sufficient diversity of new extended experts.To ensure that each specialized expert in the MoE model works as expected, we select a small amount of seed data that each expert excels to preoptimize the router.Extensive experiments with various data scales and upcycling settings demonstrate the outstanding performance and data efficiency of UpIT, as well as stable improvement in expert or data scaling.Further analysis reveals the importance of ensuring expert diversity in upcycling.
Tingfeng Hui, Zhenyu Zhang 0006, Shuohuan Wang, Yu Sun 0029, Hua Wu 0003, Sen Su
ACL (1)2
2025 Debiasing Multimodal Large Language Models via Noise-Aware Preference Optimization
abstract
Multimodal Large Language Models (MLLMs) excel in various tasks, yet often struggle with modality bias, where the model tends to rely heavily on a single modality and overlook critical information in other modalities, which leads to incorrect focus and generating irrelevant responses. In this paper, we propose using the paradigm of preference optimization to solve the modality bias problem, including RLAIF-V-Bias, a debiased preference optimization dataset, and a Noise-Aware Preference Optimization (NaPO) algorithm. Specifically, we first construct the dataset by introducing perturbations to reduce the informational content of certain modalities, compelling the model to rely on a specific modality when generating negative responses. To address the inevitable noise in automatically constructed data, we combine the noise-robust Mean Absolute Error (MAE) with the Binary Cross-Entropy (BCE) in Direct Preference Optimization (DPO) by a negative Box-Cox transformation, and dynamically adjust the algorithm’s noise robustness based on the evaluated noise levels in the data. Extensive experiments validate our approach, demonstrating not only its effectiveness in mitigating modality bias but also its significant role in minimizing hallucinations. The code and data is available at https://github.com/zhangzef/NaPO.
Zefeng Zhang 0001, Hengzhu Tang, Jiawei Sheng, Zhenyu Zhang 0006, Dawei Yin 0001, Duohe Ma, Tingwen Liu
CVPR4
2025 Mixture of Hidden-Dimensions: Not All Hidden-States' Dimensions are Needed in Transformer
abstract
Transformer models encounter inefficiency when scaling hidden dimensions due to the uniform expansion of parameters. When delving into the sparsity of hidden dimensions, we observe that only a small subset of dimensions are highly activated, where some dimensions are commonly activated across tokens, and some others uniquely activated for individual tokens. To leverage this, we propose MoHD (Mixture of Hidden Dimensions), a sparse architecture that combines shared sub-dimensions for common features and dynamically routes specialized sub-dimensions per token. To address the potential information loss from sparsity, we introduce activation scaling and group fusion mechanisms. MoHD efficiently expands hidden dimensions with minimal computational increases, outperforming vanilla Transformers in both parameter efficiency and task performance across 10 NLP tasks. MoHD achieves 1.7% higher performance with 50% fewer activatied parameters and 3.7% higher performance with 3$\times$ total parameters expansion at constant activated parameters cost. MoHD offers a new perspective for scaling the model, showcasing the potential of hidden dimension sparsity.
Junyuan Shang, Zhenyu Zhang 0006, Jiawei Sheng, Tingwen Liu, Shuohuan Wang, Yu Sun 0029, Hua Wu 0003, Haifeng Wang 0001
ICML3
2025 Multi-Branch Collaborative Learning Network for Video Quality Assessment in Industrial Video Search
abstract
Video Quality Assessment (VQA) is a crucial component of broadscale video retrieval systems. Its goal is to accurately identify various quality issues in videos, thereby encouraging the video retrieval system to prioritize high-quality videos. In large-scale industrial video retrieval systems, we formulate the characteristics of low-quality videos into four categories: visual-related low-level quality problems such as mosaics and black boxes, textual-related low-level quality problems caused by video title and Optical Character Recognition (OCR) content, as well as semantic-level frame incoherence and frame-text mismatch caused by emerging AI-generated videos. These kinds of low-quality videos, which are widely present in industrial environments, have been overlooked in academic research before, and accurately identifying them is very challenging. In this paper, we introduce a Multi-Branch Collaborative learning Network (MBCN) to tackle the above issues. We carefully design four assessment branches for MBCN to adapt to the above four kinds of issues for industrial video retrieval systems. After obtaining independent scores for each branch, we perform a weighted aggregation of the various branches to dynamically address video quality issues in different scenarios with a squeeze-and-excitation mechanism. Finally, we integrate point-wise and pair-wise optimization objectives to ensure the predicted scores are stable and fall into a reasonable range. To demonstrate the effectiveness of our proposed MBCN, we conduct extensive offline and online experiments in a world-level video search engine. The experimental results show that due to the powerful ability of MBCN to identify video quality issues, the ranking ability of the video retrieval system has been significantly improved. We also conduct a series of detailed experimental analyses to verify that all four evaluation branches play a positive role. Besides that, for emerging low-quality AI-generated videos, the recognition accuracy of MBCN also improves significantly compared to the baseline.
Hengzhu Tang, Zefeng Zhang 0001, Zhiping Li, Zhenyu Zhang 0006, Suqi Cheng, Dawei Yin 0001
KDD (1)4
2024 LEMON: Reviving Stronger and Smaller LMs from Larger LMs with Linear Parameter Fusion
abstract
Yilong Chen, Junyuan Shang, Zhenyu Zhang, Shiyao Cui, Tingwen Liu, Shuohuan Wang, Yu Sun, Hua Wu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Junyuan Shang, Zhenyu Zhang 0006, Shiyao Cui, Tingwen Liu, Shuohuan Wang, Yu Sun 0029, Hua Wu 0003
ACL (1)3
2024 NACL: A General and Effective KV Cache Eviction Framework for LLM at Inference Time
abstract
Yilong Chen, Guoxia Wang, Junyuan Shang, Shiyao Cui, Zhenyu Zhang, Tingwen Liu, Shuohuan Wang, Yu Sun, Dianhai Yu, Hua Wu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Guoxia Wang, Junyuan Shang, Shiyao Cui, Zhenyu Zhang 0006, Tingwen Liu, Shuohuan Wang, Dianhai Yu, Hua Wu 0003
ACL (1)5
2024 LoginMEA: Local-to-Global Interaction Network for Multi-Modal Entity Alignment
abstract
Multi-modal entity alignment (MMEA) aims to identify equivalent entities between two multi-modal knowledge graphs (MMKGs), whose entities can be associated with relational triples and related images. Most previous studies treat the graph structure as a special modality, and fuse different modality information with separate uni-modal encoders, neglecting valuable relational associations in modalities. Other studies refine each uni-modal information with graph structures, but may introduce unnecessary relations in specific modalities. To this end, we propose a novel local-to-global interaction network for MMEA, termed as LoginMEA. Particularly, we first fuse local multi-modal interactions to generate holistic entity semantics and then refine them with global relational interactions of entity neighbors. In this design, the uni-modal information is fused adaptively, and can be refined with relations accordingly. To enrich local interactions of multi-modal entity information, we devise modality weights and low-rank interactive fusion, allowing diverse impacts and element-level feature interactions among modalities. To capture global interactions of graph structures, we adopt relation reflection graph attention networks, which fully capture relational associations between entities. Extensive experiments demonstrate superior results of our method over 5 cross-KG or bilingual benchmark datasets, indicating the effectiveness of capturing local and global interactions.
Taoyu Su, Xinghua Zhang 0001, Jiawei Sheng, Zhenyu Zhang 0006, Tingwen Liu
ECAI4
2024 DHA: Learning Decoupled-Head Attention from Transformer Checkpoints via Adaptive Heads Fusion
abstract
Large language models (LLMs) with billions of parameters demonstrate impressive performance. However, the widely used Multi-Head Attention (MHA) in LLMs incurs substantial computational and memory costs during inference. While some efforts have optimized attention mechanisms by pruning heads or sharing parameters among heads, these methods often lead to performance degradation or necessitate substantial continued pre-training costs to restore performance. Based on the analysis of attention redundancy, we design a Decoupled-Head Attention (DHA) mechanism. DHA adaptively configures group sharing for key heads and value heads across various layers, achieving a better balance between performance and efficiency. Inspired by the observation of clustering similar heads, we propose to progressively transform the MHA checkpoint into the DHA model through linear fusion of similar head parameters step by step, retaining the parametric knowledge of the MHA checkpoint. We construct DHA models by transforming various scales of MHA checkpoints given target head budgets. Our experiments show that DHA remarkably requires a mere 0.25\% of the original model's pre-training budgets to achieve 96.1\% of performance while saving 75\% of KV cache. Compared to Group-Query Attention (GQA), DHA achieves a 5$\times$ training acceleration, a maximum of 13.93\% performance improvement under 0.01\% pre-training budget, and 5\% relative improvement under 0.05\% pre-training budget.
Linhao Zhang, Junyuan Shang, Zhenyu Zhang 0006, Tingwen Liu, Shuohuan Wang
NeurIPS4
2023 ERNIE-ViLG 2.0: Improving Text-to-Image Diffusion Model with Knowledge-Enhanced Mixture-of-Denoising-Experts
abstract
Recent progress in diffusion models has revolutionized the popular technology of text-to-image generation. While existing approaches could produce photorealistic high-resolution images with text conditions, there are still several open problems to be solved, which limits the further improvement of image fidelity and text relevancy. In this paper, we propose ERNIE-ViLG 2.0, a large-scale Chinese text-to-image diffusion model, to progressively upgrade the quality of generated images by: (1) incorporating fine-grained textual and visual knowledge of key elements in the scene, and (2) utilizing different denoising experts at different denoising stages. With the proposed mechanisms, ERNIE-ViLG 2.01not only achieves a new state-of-the-art on MS-COCO with zero-shot FID-30k score of 6.75, but also significantly outperforms recent models in terms of image fidelity and image-text alignment, with side-by-side human evaluation on the bilingual prompt set ViLG-300.
Zhida Feng, Zhenyu Zhang 0006, Yewei Fang, Lanxin Li, Xuyi Chen, Jiaxiang Liu 0004, Weichong Yin, Shikun Feng, Yu Sun 0004, Li Chen 0011, Hao Tian 0005, Hua Wu 0003, Haifeng Wang 0001
CVPR2
2023 Enhancing Table Retrieval with Dual Graph Representations
Tianyun Liu, Xinghua Zhang 0001, Zhenyu Zhang 0006, Quangang Li, Tingwen Liu
ECML/PKDD (4)3
2023 Learning Structural Co-occurrences for Structured Web Data Extraction in Low-Resource Settings
abstract
Extracting structured information from all manner of webpages is an important problem with the potential to automate many real-world applications. Recent work has shown the effectiveness of leveraging DOM trees and pre-trained language models to describe and encode webpages. However, they typically optimize the model to learn the semantic co-occurrence of elements and labels in the same webpage, thus their effectiveness depends on sufficient labeled data, which is labor-intensive. In this paper, we further observe structural co-occurrences in different webpages of the same website: the same position in the DOM tree usually plays the same semantic role, and the DOM nodes in this position also share similar surface forms. Motivated by this, we propose a novel method, Structor, to effectively incorporate the structural co-occurrences over DOM tree and surface form into pre-trained language models. Such structural co-occurrences help the model learn the task better under low-resource settings, and we study two challenging experimental scenarios: website-level low-resource setting and webpage-level low-resource setting, to evaluate our approach. Extensive experiments on the public SWDE dataset show that Structor significantly outperforms the state-of-the-art models in both settings, and even achieves three times the performance of the strong baseline model in the case of extreme lack of training data.
Zhenyu Zhang 0006, Bowen Yu 0002, Tingwen Liu, Tianyun Liu, Li Guo 0001
WWW1
2022 Enhancing Chinese Pre-trained Language Model via Heterogeneous Linguistics Graph
abstract
Yanzeng Li, Jiangxia Cao, Xin Cong, Zhenyu Zhang, Bowen Yu, Hongsong Zhu, Tingwen Liu. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Yanzeng Li, Jiangxia Cao, Xin Cong, Zhenyu Zhang 0006, Bowen Yu 0002, Hongsong Zhu, Tingwen Liu
ACL (1)4
2022 Layout-Aware Information Extraction for Document-Grounded Dialogue: Dataset, Method and Demonstration
abstract
Building document-grounded dialogue systems have received growing interest as documents convey a wealth of human knowledge and commonly exist in enterprises. Wherein, how to comprehend and retrieve information from documents is a challenging research problem. Previous work ignores the visual property of documents and treats them as plain text, resulting in incomplete modality. In this paper, we propose a Layout-aware document-level Information Extraction dataset, LIE, to facilitate the study of extracting both structural and semantic knowledge from visually rich documents (VRDs), so as to generate accurate responses in dialogue systems. LIE contains 62k annotations of three extraction tasks from 4,061 pages in product and official documents, becoming the largest VRD-based information extraction dataset to the best of our knowledge. We also develop benchmark methods that extend the token-based language model to consider layout features like humans. Empirical results show that layout is critical for VRD-based extraction, and system demonstration also verifies that the extracted knowledge can help locate the answers that users care about.
Zhenyu Zhang 0006, Bowen Yu 0002, Haiyang Yu 0003, Tingwen Liu, Cheng Fu 0003, Chengguang Tang, Jian Sun 0021, Yongbin Li 0001
ACM Multimedia1
2021 Improving Distantly-Supervised Named Entity Recognition with Self-Collaborative Denoising Learning
abstract
Distantly supervised named entity recognition (DS-NER) efficiently reduces labor costs but meanwhile intrinsically suffers from the label noise due to the strong assumption of distant supervision.Typically, the wrongly labeled instances comprise numbers of incomplete and inaccurate annotation noise, while most prior denoising works are only concerned with one kind of noise and fail to fully explore useful information in the whole training set.To address this issue, we propose a robust learning paradigm named Self-Collaborative Denoising Learning (SCDL), which jointly trains two teacher-student networks in a mutuallybeneficial manner to iteratively perform noisy label refinery.Each network is designed to exploit reliable labels via self denoising, and two networks communicate with each other to explore unreliable annotations by collaborative denoising.Extensive experimental results on five real-world datasets demonstrate that SCDL is superior to state-of-the-art DS-NER denoising methods 1 .
Xinghua Zhang 0001, Bowen Yu 0002, Tingwen Liu, Zhenyu Zhang 0006, Jiawei Sheng, Mengge Xue
EMNLP (1)4
2021 Multi-Granularity Heterogeneous Graph for Document-Level Relation Extraction
abstract
Reading text to extract relational facts has been a long-standing goal in natural language processing. It becomes especially challenging when the extraction scope is extended to document level, where multiple entities in a document generally exhibit complex intra- and inter-sentence relations. In this paper, we propose a novel Multi-granularity Heterogeneous Graph (MHG) to tackle this challenge. Specifically, we define four types of nodes with different granularities and eight types of edges based on heuristic rules, entrusting the MHG two major advantages. On the one hand, it connects any two entities with a short path in the graph to better handle the complex inter-sentence interactions between entities. On the other hand, it enables rich interactions among nodes with different granularities to promote accurate multi-hop reasoning. Experimental results on the largest document-level relation extraction dataset suggest that the proposed model achieves new state-of-the-art performance.
Hengzhu Tang, Yanan Cao 0001, Zhenyu Zhang 0006, Ruipeng Jia, Fang Fang 0009, Shi Wang 0002
ICASSP3
2021 NA-Aware Machine Reading Comprehension for Document-Level Relation Extraction
Zhenyu Zhang 0006, Bowen Yu 0002, Xiaobo Shu, Tingwen Liu
ECML/PKDD (3)1
2021 Semi-Open Information Extraction
abstract
Open Information Extraction (OIE), the task aimed at discovering all textual facts organized in the form of (subject, predicate, object) found within a sentence, has gained much attention recently. However, in some knowledge-driven applications such as question answering, we often have a target entity and hope to obtain its structured factual knowledge for better understanding, instead of extracting all possible facts aimlessly from the corpus. In this paper, we define a new task, namely Semi-Open Information Extraction (SOIE), to address this need. The goal of SOIE is to discover domain-independent facts towards a particular entity from general and diverse web text. To facilitate research on this new task, we propose a large-scale human-annotated benchmark called SOIED, consisting of 61,984 facts for 8,013 subject entities annotated on 24,000 Chinese sentences collected from the web search engine.
Bowen Yu 0002, Zhenyu Zhang 0006, Jiawei Sheng, Tingwen Liu, Bin Wang 0004
WWW2
2020 Distilling Knowledge from Well-Informed Soft Labels for Neural Relation Extraction
abstract
Extracting relations from plain text is an important task with wide application. Most existing methods formulate it as a supervised problem and utilize one-hot hard labels as the sole target in training, neglecting the rich semantic information among relations. In this paper, we aim to explore the supervision with soft labels in relation extraction, which makes it possible to integrate prior knowledge. Specifically, a bipartite graph is first devised to discover type constraints between entities and relations based on the entire corpus. Then, we combine such type constraints with neural networks to achieve a knowledgeable model. Furthermore, this model is regarded as teacher to generate well-informed soft labels and guide the optimization of a student network via knowledge distillation. Besides, a multi-aspect attention mechanism is introduced to help student mine latent information from text. In this way, the enhanced student inherits the dark knowledge (e.g., type constraints and relevance among relations) from teacher, and directly serves the testing scenarios without any extra constraints. We conduct extensive experiments on the TACRED and SemEval datasets, the experimental results justify the effectiveness of our approach.
Zhenyu Zhang 0006, Xiaobo Shu, Bowen Yu 0002, Tingwen Liu, Jiapeng Zhao, Quangang Li, Li Guo 0001
AAAI1
2020 Learning to Prune Dependency Trees with Rethinking for Neural Relation Extraction
abstract
Dependency trees have been shown to be effective in capturing long-range relations between target entities.Nevertheless, how to selectively emphasize target-relevant information and remove irrelevant content from the tree is still an open problem.Existing approaches employing predefined rules to eliminate noise may not always yield optimal results due to the complexity and variability of natural language.In this paper, we present a novel architecture named Dynamically Pruned Graph Convolutional Network (DP-GCN), which learns to prune the dependency tree with rethinking in an end-to-end scheme.In each layer of DP-GCN, we employ a selection module to concentrate on nodes expressing the target relation by a set of binary gates and then augment the pruned tree with a pruned semantic graph to ensure the connectivity.After that, we introduce a rethinking mechanism to guide and refine the pruning operation by feeding back the high-level learned features repeatedly.Extensive experimental results demonstrate that our model achieves impressive performance compared to strong competitors.
Bowen Yu 0002, Mengge Xue, Zhenyu Zhang 0006, Tingwen Liu, Bin Wang 0004
COLING3
2020 Document-level Relation Extraction with Dual-tier Heterogeneous Graph
abstract
Document-level relation extraction (RE)poses new challenges over its sentence-level counterpart since it requires an adequate comprehension of the whole document and the multi-hop reasoning ability across multiple sentences to reach the final result.In this paper, we propose a novel graphbased model with Dual-tier Heterogeneous Graph (DHG) for document-level RE.In particular, DHG is composed of a structure modeling layer followed by a relation reasoning layer.The major advantage is that it is capable of not only capturing both the sequential and structural information of documents but also mixing them together to benefit for multi-hop reasoning and final decisionmaking.Furthermore, we employ Graph Neural Networks (GNNs) based message propagation strategy to accumulate information on DHG.Experimental results demonstrate that the proposed method achieves state-of-the-art performance on two widely used datasets, and further analyses suggest that all the modules in our model are indispensable for document-level RE.
Zhenyu Zhang 0006, Bowen Yu 0002, Xiaobo Shu, Tingwen Liu, Hengzhu Tang, Li Guo 0001
COLING1
2020 Joint Extraction of Entities and Relations Based on a Novel Decomposition Strategy
abstract
Joint extraction of entities and relations aims to detect entity pairs along with their relations using a single model. Prior work typically solves this task in the extract-then-classify or unified labeling manner. However, these methods either suffer from the redundant entity pairs, or ignore the important inner structure in the process of extracting entities and relations. To address these limitations, in this paper, we first decompose the joint extraction task into two interrelated subtasks, namely HE extraction and TER extraction. The former subtask is to distinguish all head-entities that may be involved with target relations, and the latter is to identify corresponding tail-entities and relations for each extracted head-entity. Next, these two subtasks are further deconstructed into several sequence labeling problems based on our proposed span-based tagging scheme, which are conveniently solved by a hierarchical boundary tagger and a multi-span decoding algorithm. Owing to the reasonable decomposition strategy, our model can fully capture the semantic interdependency between different steps, as well as reduce noise from irrelevant entity pairs. Experimental results show that our method outperforms previous work by 5.2%, 5.9% and 21.5% (F1 score), achieving a new state-of-the-art on three public datasets.
Bowen Yu 0002, Zhenyu Zhang 0006, Xiaobo Shu, Tingwen Liu, Bin Wang 0004, Sujian Li
ECAI2
2020 Coarse-to-Fine Pre-training for Named Entity Recognition
abstract
More recently, Named Entity Recognition has achieved great advances aided by pre-training approaches such as BERT.However, current pre-training techniques focus on building language modeling objectives to learn a general representation, ignoring the named entityrelated knowledge.To this end, we propose a NER-specific pre-training framework to inject coarse-to-fine automatically mined entity knowledge into pre-trained models.Specifically, we first warm-up the model via an entity span identification task by training it with Wikipedia anchors, which can be deemed as general-typed entities.Then we leverage the gazetteer-based distant supervision strategy to train the model extract coarse-grained typed entities.Finally, we devise a self-supervised auxiliary task to mine the fine-grained named entity knowledge via clustering.Empirical studies on three public NER datasets demonstrate that our framework achieves significant improvements against several pre-trained baselines, establishing the new state-of-the-art performance on three benchmarks.Besides, we show that our framework gains promising results without using human-labeled training data, demonstrating its effectiveness in labelfew and low-resource scenarios.1
Mengge Xue, Bowen Yu 0002, Zhenyu Zhang 0006, Tingwen Liu, Bin Wang 0004
EMNLP (1)3
2020 Joint Entity Linking and Relation Extraction with Neural Networks for Knowledge Base Population
abstract
Relation extraction and entity linking are two fundamental procedures to extend knowledge bases. Most existing methods typically treat them separately and ignore the semantic relevance between entities and relations. In this paper, we pioneer a general joint learning framework for relation extraction and entity linking, which allows these two tasks boost each other. Based on the framework, a demonstration model is proposed with neural networks. We conduct experiments on variants of a standard benchmark dataset (NYT-10) to verify the effectiveness of our approach. Experimental results show that our approach significantly outperforms traditional separate methods without reducing efficiency, especially on datasets with many ambiguous entity mentions. Furthermore, various mainstream methods for relation extraction and entity linking can be easily integrated into our loosely-coupled framework due to its flexible architecture.
Zhenyu Zhang 0006, Xiaobo Sind, Tingwen Liu, Zheng Fang 0002, Quangang Li
IJCNN1
2020 DRG2vec: Learning Word Representations from Definition Relational Graph
abstract
Even with larger and larger text data available, encoding linguistic knowledge into word embeddings directly from corpora is still difficult. The most intuitive way to incorporate knowledge into word embeddings is to use external resources, such as the largest datasource for describing words - natural language dictionaries. However, previous methods usually neglect the recursive nature of dictionaries and cannot capture the high-order relation. In this paper, we present DRG2vec, a novel and efficient approach for learning word representations, which exploits the inherent recursiveness of dictionaries by modeling the whole dictionary as a homogeneous graph based on the co-occurrence of entry and word in the definition. Moreover, a tailor-made sampling strategy is introduced to generate word sequences from the definition relational graph, and then, the generated sequences are fed to the Skip-gram model with semantic negative sampling for word representation learning. Extensive experiments on sixteen benchmark datasets show that leveraging the recursive dictionary graph indeed achieves better performance than other state-of-the-art methods, especially exhibits a weighted average improvement of 8.6% in the word similarity task.
Xiaobo Shu, Bowen Yu 0002, Zhenyu Zhang 0006, Tingwen Liu
IJCNN3
2020 BiG-Transformer: Integrating Hierarchical Features for Transformer via Bipartite Graph
abstract
Self-attention based models like Transformer have achieved great success on kinds of Natural Language Processing tasks. However, the traditional fixed fully-connected structure faces many challenges in practice, such as computing redundancy, fixed granularity, and inexplicable. In this paper, we present BiG-Transformer, which employs attention with bipartite-graph structure to replace the fully-connected self-attention mechanism in Transformer. Specifically, two parts of the graph are designed for integrating hierarchical semantic information, and two types of connection are proposed to fuse information from different positions. Experiments on four tasks show the BiG-Transformer achieves better performance compared to Transformer liked models and Recurrent Neural Networks.
Xiaobo Shu, Mengge Xue, Yanzeng Li, Zhenyu Zhang 0006, Tingwen Liu
IJCNN4
2020 Strong Baselines for Author Name Disambiguation with and Without Neural Networks
Zhenyu Zhang 0006, Bowen Yu 0002, Tingwen Liu, Dong Wang 0029
PAKDD (1)1
2020 HIN: Hierarchical Inference Network for Document-Level Relation Extraction
Hengzhu Tang, Yanan Cao 0001, Zhenyu Zhang 0006, Jiangxia Cao, Fang Fang 0009, Shi Wang 0002, Pengfei Yin
PAKDD (1)3
2020 SLGAT: Soft Labels Guided Graph Attention Networks
Zhenyu Zhang 0006, Tingwen Liu, Li Guo 0001
PAKDD (1)2
2020 Fine-Grained Semantics-Aware Heterogeneous Graph Neural Networks
Zhenyu Zhang 0006, Tingwen Liu, Li Guo 0001
WISE (1)2
2020 High Quality Candidate Generation and Sequential Graph Attention Network for Entity Linking
abstract
Entity Linking (EL) is a task for mapping mentions in text to corresponding entities in knowledge base (KB). This task usually includes candidate generation (CG) and entity disambiguation (ED) stages. Recent EL systems based on neural network models have achieved good performance, but they still face two challenges: (i) Previous studies evaluate their models without considering the differences between candidate entities. In fact, the quality (gold recall in particular) of candidate sets has an effect on the EL results. So, how to promote the quality of candidates needs more attention. (ii) In order to utilize the topical coherence among the referred entities, many graph and sequence models are proposed for collective ED. However, graph-based models treat all candidate entities equally which may introduce much noise information. On the contrary, sequence models can only observe previous referred entities, ignoring the relevance between the current mention and its subsequent entities. To address the first problem, we propose a multi-strategy based CG method to generate high recall candidate sets. For the second problem, we design a Sequential Graph Attention Network (SeqGAT) which combines the advantages of graph and sequence methods. In our model, mentions are dealt with in a sequence manner. Given the current mention, SeqGAT dynamically encodes both its previous referred entities and subsequent ones, and assign different importance to these entities. In this way, it not only makes full use of the topical consistency, but also reduce noise interference. We conduct experiments on different types of datasets and compare our method with previous EL system on the open evaluation platform. The comparison results show that our model achieves significant improvements over the state-of-the-art methods.
Zheng Fang 0002, Yanan Cao 0001, Zhenyu Zhang 0006, Yanbing Liu 0007, Shi Wang 0002
WWW4
2019 Beyond Word Attention: Using Segment Attention in Neural Relation Extraction
abstract
Relation extraction studies the issue of predicting semantic relations between pairs of entities in sentences. Attention mechanisms are often used in this task to alleviate the inner-sentence noise by performing soft selections of words independently. Based on the observation that information pertinent to relations is usually contained within segments (continuous words in a sentence), it is possible to make use of this phenomenon for better extraction. In this paper, we aim to incorporate such segment information into neural relation extractor. Our approach views the attention mechanism as linear-chain conditional random fields over a set of latent variables whose edges encode the desired structure, and regards attention weight as the marginal distribution of each word being selected as a part of the relational expression. Experimental results show that our method can attend to continuous relational expressions without explicit annotations, and achieve the state-of-the-art performance on the large-scale TACRED dataset.
Bowen Yu 0002, Zhenyu Zhang 0006, Tingwen Liu, Bin Wang 0004, Sujian Li, Quangang Li
IJCAI2
2019 ICNet: Incorporating Indicator Words and Contexts to Identify Functional Description Information
abstract
Functional description information refers to the texts that describe the functionality or performance characteristics of a certain object. This type of information is of great potential value for the field of intelligence discovery. Thus automatically and accurately identifying this information from large amounts of texts on the web is very important. In this paper we reduce the functional description problem to a binary classification task deciding whether the input sentence is a functional description sentence or not. However, there exist lots of comment texts in the web data, which are semantically very similar to description texts, making our task quite difficult. Also, existing methods only provide general sentence representation models, which can't lead to targeted ways to solve our problem. Therefore, to address the problem, we not only exploit contexts, like many other previous work did, but also introduce indicator word information to learn rich representations. And in order to incorporate them both, we propose two models, namely ICNet(multi-tasks) and ICNet(ensemble). ICNet(multitasks) exploits them jointly in a integrated process of learning representations, while ICNet(ensemble) exploits them by two respective but concatenated sub-models. Experimental results on the collected real-world dataset indicate that both ICNet(multitasks) and ICNet(ensemble) achieve higher F1 scores compared with FaxtText, CNN, RNN, LSTM and Bi-LSTM, QuickThought models on this task.
Qu Liu, Zhenyu Zhang 0006, Yanzeng Li, Tingwen Liu, Diying Li, Jinqiao Shi
IJCNN2
2019 Joint Entity Linking with Deep Reinforcement Learning
abstract
Entity linking is the task of aligning mentions to corresponding entities in a given knowledge base. Previous studies have highlighted the necessity for entity linking systems to capture the global coherence. However, there are two common weaknesses in previous global models. First, most of them calculate the pairwise scores between all candidate entities and select the most relevant group of entities as the final result. In this process, the consistency among wrong entities as well as that among right ones are involved, which may introduce noise data and increase the model complexity. Second, the cues of previously disambiguated entities, which could contribute to the disambiguation of the subsequent mentions, are usually ignored by previous models. To address these problems, we convert the global linking into a sequence decision problem and propose a reinforcement learning model which makes decisions from a global perspective. Our model makes full use of the previous referred entities and explores the long-term influence of current selection on subsequent decisions. We conduct experiments on different types of datasets, the results show that our model outperforms state-of-the-art systems and has better generalization performance.
Zheng Fang 0002, Yanan Cao 0001, Qian Li 0003, Zhenyu Zhang 0006, Yanbing Liu 0007
WWW5