Kunyuan Pang

dblp:168/0868 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
7since 2021 · last 2023
0000-0002-9477-4458ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Information extraction and text analysis · 100%

Topics — the 2 heaviest of 2, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Information extraction and text analysis
entity typing
0.612022
Divide and Denoise: Learning from Noisy Labels in Fine-Grained Entity Typing with Cluster-Wise Loss Correction · ACL (1) 2022
Natural language and speech › Information extraction and text analysis › entity typing
fine-grained entity typing
0.612022
Divide and Denoise: Learning from Noisy Labels in Fine-Grained Entity Typing with Cluster-Wise Loss Correction · ACL (1) 2022

Methods — techniques the papers use, named apart from their topics

loss correction · 0.6feature extraction · 0.6clustering · 0.6
YearPublicationVenuePosition
2023 Distinguishing Sensitive and Insensitive Options for the Winograd Schema Challenge
Dong Li 0048, Pancheng Wang, Liangliang He, Kunyuan Pang, Shasha Li 0001, Jintao Tang, Ting Wang 0009
DASFAA (3)4
2022 Divide and Denoise: Learning from Noisy Labels in Fine-Grained Entity Typing with Cluster-Wise Loss Correction
abstract
Fine-grained Entity Typing (FET) has made great progress based on distant supervision but still suffers from label noise.Existing FET noise learning methods rely on prediction distributions in an instance-independent manner, which causes the problem of confirmation bias.In this work, we propose a clustering-based loss correction framework named Feature Cluster Loss Correction (FCLC), to address these two problems.FCLC first train a coarse backbone model as a feature extractor and noise estimator.Loss correction is then applied to each feature cluster, learning directly from the noisy labels.Experimental results on three public datasets show that FCLC achieves the best performance over existing competitive systems.Auxiliary experiments further demonstrate that FCLC is stable to hyperparameters and it does help mitigate confirmation bias.We also find that in the extreme case of no clean data, the FCLC framework still achieves competitive performance.
Kunyuan Pang, Jie Zhou 0013, Ting Wang 0009
ACL (1)1
2022 Multi-Document Scientific Summarization from a Knowledge Graph-Centric View
abstract
Multi-Document Scientific Summarization (MDSS) aims to produce coherent and concise summaries for clusters of topic-relevant scientific papers. This task requires precise understanding of paper content and accurate modeling of cross-paper relationships. Knowledge graphs convey compact and interpretable structured information for documents, which makes them ideal for content modeling and relationship modeling. In this paper, we present KGSum, an MDSS model centred on knowledge graphs during both the encoding and decoding process. Specifically, in the encoding process, two graph-based modules are proposed to incorporate knowledge graph information into paper encoding, while in the decoding process, we propose a two-stage decoder by first generating knowledge graph information of summary in the form of descriptive sentences, followed by generating the final summary. Empirical results show that the proposed architecture brings substantial improvements over baselines on the Multi-Xscience dataset.
Pancheng Wang, Shasha Li 0001, Kunyuan Pang, Liangliang He, Dong Li 0048, Jintao Tang, Ting Wang 0009
COLING3
2022 Multi-source Representation Enhancement for Wikipedia-style Entity Annotation
abstract
Entity annotation in Wikipedia (officially named wikilinks) greatly benefits human end-users. Human editors are required to select all mentions that are most helpful to human end-users and link each mention to a Wikipedia page. We aim to design an automatic system to generate Wikipedia-style entity annotation for any plain text. However, existing research either rely heavily on mention-entity map or are restricted to named entities only. Besides, they neglect to select the appropriate mentions as Wikipedia requires. As a result, they leave out some necessary annotation and introduce excessive distracting annotation. Existing benchmarks also skirt around the coverage and selection issues. We propose a new task called Mention Detection and Se-lection for entity annotation, along with a new benchmark, WikiC, to better reflect annotation quality. The task is coined centering mentions specific to each position in high-quality human-annotated examples. We also proposed a new framework, DrWiki, to fulfill the task. We adopt a deep pre-trained span selection model inferring directly from plain text via tokens' context embedding. It can cover all possible spans and avoid limiting to mention-entity maps. In addition, information of both inarguable mention-entity pairs, and mention repeat has been introduced as token-wise representation enhancement by FLAT attention and repeat embedding respectively. Empirical results on WikiC show that, compared with often adopted and state-of-the-art Entity Linking and Entity Recognition methods, our method achieves improvement to previous methods in overall performance. Additional experiments show that DrWiki gains improvement even with a low-coverage mention-entity map.
Kunyuan Pang, Shasha Li 0001, Jintao Tang, Ting Wang 0009
IJCNN1
2022 BERT-SMAP: Paying attention to Essential Terms in passage ranking beyond BERT
Dengwen Lin, Jintao Tang, Xinyi Li 0001, Kunyuan Pang, Shasha Li 0001, Ting Wang 0009
Inf. Process. Manag.4
2021 Show, Rethink, And Tell: Image Caption Generation With Hierarchical Topic Cues
abstract
Current state-of-the-art approaches for image captioning mainly apply the encoder-decoder framework with attention mechanisms, most of which ignore interactions between different types of image features and perform attention operations only once per word. The mentioned problems limit the captioning model’s capability to capture sufficient information to generate high-quality captions. By contrast, humans often rethink to polish up descriptions by re-focusing on more correct and important information, which is hard to capture at first glance. In this paper, we introduce a novel topic-guided captioning model to imitate such a human’s rethinking process by modeling interactions between visual and hierarchical semantic features of topics. To the best of our knowledge, we are the first to effectively consider hierarchical semantic features as guidance to facilitate visual attention, achieving human-like rethinking for captioning. Extensive experiments on the MS COCO dataset show that our proposed model achieves superior performance over state-of-the-art methods.
Songxian Xie, Xinyi Li 0001, Jintao Tang, Kunyuan Pang, Shasha Li 0001, Ting Wang 0009
ICME5
2021 Constructing an Educational Knowledge Graph with Concepts Linked to Wikipedia
Fu-Rong Dang, Jintao Tang, Kunyuan Pang, Ting Wang 0009, Shasha Li 0001, Xiao Li 0039
J. Comput. Sci. Technol.3
2018 Which Embedding Level is Better for Semantic Representation? An Empirical Research on Chinese Phrases
Kunyuan Pang, Jintao Tang, Ting Wang 0009
NLPCC (2)1