Kang Zhou 0002

dblp:50/10844-2 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
7since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Information extraction and text analysis · 100%
Databases, data mining, and information retrieval
2 papers
Knowledge graphs · 51% Machine learning and data management · 34% Data mining · 15%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Information extraction and text analysis › relation extraction
open relation extraction
1.522025
Towards a More Generalized Approach in Open Relation Extraction · ACL (1) 2025
Improving Unsupervised Relation Extraction by Augmenting Diverse Sentence Pairs · EMNLP 2023
Natural language and speech › Information extraction and text analysis
relation extraction
1.322023
Improving Unsupervised Relation Extraction by Augmenting Diverse Sentence Pairs · EMNLP 2023
Improving Distantly Supervised Relation Extraction by Natural Language Inference · AAAI 2023
Knowledge graphs
relation extraction
0.912025
Towards a More Generalized Approach in Open Relation Extraction · ACL (1) 2025
Natural language and speech › Information extraction and text analysis › relation extraction
distant supervision
0.712023
Improving Distantly Supervised Relation Extraction by Natural Language Inference · AAAI 2023
Natural language and speech › Information extraction and text analysis › named entity recognition
distantly supervised NER
0.612022
Distantly Supervised Named Entity Recognition via Confidence-Based Multi-Class Positive and Unlabeled Learning · ACL (1) 2022
Natural language and speech › Information extraction and text analysis
named entity recognition
0.612022
Distantly Supervised Named Entity Recognition via Confidence-Based Multi-Class Positive and Unlabeled Learning · ACL (1) 2022
Machine learning and data management › weak supervision
positive-unlabeled learning
0.612022
Distantly Supervised Named Entity Recognition via Confidence-Based Multi-Class Positive and Unlabeled Learning · ACL (1) 2022
Data mining
clustering
0.312025
Towards a More Generalized Approach in Open Relation Extraction · ACL (1) 2025

Methods — techniques the papers use, named apart from their topics

relation classification · 1.7clustering · 1.7multi-class positive and unlabeled learning · 1.1confidence-based risk estimation · 1.1relation verbalization · 0.7natural language inference · 0.7margin loss · 0.7data consolidation · 0.7data augmentation · 0.7contrastive learning · 0.7
YearPublicationVenuePosition
2025 Towards a More Generalized Approach in Open Relation Extraction
abstract
Open Relation Extraction (OpenRE) seeks to identify and extract novel relational facts between named entities from unlabeled data without pre-defined relation schemas. Traditional OpenRE methods typically assume that the unlabeled data consists solely of novel relations or is pre-divided into known and novel instances. However, in real-world scenarios, novel relations are arbitrarily distributed. In this paper, we propose a generalized OpenRE setting that considers unlabeled data as a mixture of both known and novel instances. To address this, we propose MixORE, a two-phase framework that integrates relation classification and clustering to jointly learn known and novel relations. Experiments on three benchmark datasets demonstrate that MixORE consistently outperforms competitive baselines in known relation classification and novel relation clustering. Our findings contribute to the advancement of generalized OpenRE research and real-world applications.
Yuepei Li, Qiao Qiao, Kang Zhou 0002, Qi Li 0012
ACL (1)4
2025 Re-Examine Distantly Supervised NER: A New Benchmark and a Simple Approach
abstract
Distantly-Supervised Named Entity Recognition (DS-NER) uses knowledge bases or dictionaries for annotations, reducing manual efforts but rely on large human labeled validation set. In this paper, we introduce a real-life DS-NER dataset, QTL, where the training data is annotated using domain dictionaries and the test data is annotated by domain experts. This dataset has a small validation set, reflecting real-life scenarios. Existing DS-NER approaches fail when applied to QTL, which motivate us to re-examine existing DS-NER approaches. We found that many of them rely on large validation sets and some used test set for tuning inappropriately. To solve this issue, we proposed a new approach, token-level Curriculum-based Positive-Unlabeled Learning (CuPUL), which uses curriculum learning to order training samples from easy to hard. This method stabilizes training, making it robust and effective on small validation sets. CuPUL also addresses false negative issues using the Positive-Unlabeled learning paradigm, demonstrating improved performance in real-life applications.
Yuepei Li, Kang Zhou 0002, Qiao Qiao, Qi Li 0012
COLING2
2025 Bridge Structural Knowledge and Pre-trained Language Models for Knowledge Graph Completion
Qiao Qiao, Yuepei Li, Kang Zhou 0002, Qi Li 0012
PAKDD (3)4
2023 Improving Distantly Supervised Relation Extraction by Natural Language Inference
abstract
To reduce human annotations for relation extraction (RE) tasks, distantly supervised approaches have been proposed, while struggling with low performance. In this work, we propose a novel DSRE-NLI framework, which considers both distant supervision from existing knowledge bases and indirect supervision from pretrained language models for other tasks. DSRE-NLI energizes an off-the-shelf natural language inference (NLI) engine with a semi-automatic relation verbalization (SARV) mechanism to provide indirect supervision and further consolidates the distant annotations to benefit multi-classification RE models. The NLI-based indirect supervision acquires only one relation verbalization template from humans as a semantically general template for each relationship, and then the template set is enriched by high-quality textual patterns automatically mined from the distantly annotated corpus. With two simple and effective data consolidation strategies, the quality of training data is substantially improved. Extensive experiments demonstrate that the proposed framework significantly improves the SOTA performance (up to 7.73% of F1) on distantly supervised RE benchmark datasets. Our code is available at https://github.com/kangISU/DSRE-NLI.
Kang Zhou 0002, Qiao Qiao, Yuepei Li, Qi Li 0012
AAAI1
2023 Improving Unsupervised Relation Extraction by Augmenting Diverse Sentence Pairs
abstract
Unsupervised relation extraction (URE) aims to extract relations between named entities from raw text without requiring manual annotations or pre-existing knowledge bases.In recent studies of URE, researchers put a notable emphasis on contrastive learning strategies for acquiring relation representations.However, these studies often overlook two important aspects: the inclusion of diverse positive pairs for contrastive learning and the exploration of appropriate loss functions.In this paper, we propose AugURE with both within-sentence pairs augmentation and augmentation through crosssentence pairs extraction to increase the diversity of positive pairs and strengthen the discriminative power of contrastive learning.We also identify the limitation of noise-contrastive estimation (NCE) loss for relation representation learning and propose to apply margin loss for sentence pairs.Experiments on NYT-FB and TACRED datasets demonstrate that the proposed relation representation learning and a simple K-Means clustering achieves state-ofthe-art performance.Source code is available 1 .
Kang Zhou 0002, Qiao Qiao, Yuepei Li, Qi Li 0012
EMNLP2
2023 Relation-Aware Network with Attention-Based Loss for Few-Shot Knowledge Graph Completion
Qiao Qiao, Yuepei Li, Kang Zhou 0002, Qi Li 0012
PAKDD (3)3
2022 Distantly Supervised Named Entity Recognition via Confidence-Based Multi-Class Positive and Unlabeled Learning
abstract
In this paper, we study the named entity recognition (NER) problem under distant supervision.Due to the incompleteness of the external dictionaries and/or knowledge bases, such distantly annotated training data usually suffer from a high false negative rate.To this end, we formulate the Distantly Supervised NER (DS-NER) problem via Multi-class Positive and Unlabeled (MPU) learning and propose a theoretically and practically novel CONFidence-based MPU (Conf-MPU) approach.To handle the incomplete annotations, Conf-MPU consists of two steps.First, a confidence score is estimated for each token of being an entity token.Then, the proposed Conf-MPU risk estimation is applied to train a multi-class classifier for the NER task.Thorough experiments on two benchmark datasets labeled by various external knowledge demonstrate the superiority of the proposed Conf-MPU over existing DS-NER methods.Our code is available at Github 1 .
Kang Zhou 0002, Yuepei Li, Qi Li 0012
ACL (1)1
2020 Fine-Grained Named Entity Recognition with Distant Supervision in COVID-19 Literature
abstract
Biomedical named entity recognition (BioNER) is a fundamental step for mining COVID-19 literature. Existing BioNER datasets cover a few common coarse-grained entity types (e.g., genes, chemicals, and diseases), which cannot be used to recognize highly domain-specific entity types (e.g., animal models of diseases) or emerging ones (e.g., coronaviruses) for COVID-19 studies. We present CORD-NER, a fine-grained named entity recognized dataset of COVID-19 literature (up until May 19, 2020). CORD-NER contains over 12 million sentences annotated via distant supervision. Also included in CORD-NER are 2,000 manually-curated sentences as a test set for performance evaluation. CORD-NER covers 75 fine-grained entity types. In addition to the common biomedical entity types, it covers new entity types specifically related to COVID-19 studies, such as coronaviruses, viral proteins, evolution, and immune responses. The dictionaries of these fine-grained entity types are collected from existing knowledge bases and human-input seed sets. We further present DISTNER, a distantly supervised NER model that relies on a massive unlabeled corpus and a collection of dictionaries to annotate the COVID-19 corpus. DISTNER provides a benchmark performance on the CORD-NER test set for future research.
Xuan Wang 0008, Xiangchen Song, Bangzheng Li, Kang Zhou 0002, Qi Li 0012, Jiawei Han 0001
BIBM4