Yu Sun 0031

dblp:62/3689-31 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Information extraction and text analysis · 54% Efficient and distributed learning · 22% Language models and text generation · 19%
Databases, data mining, and information retrieval
1 paper
Data integration and cleaning · 100%

Topics — the 5 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data integration and cleaning
data quality
0.912025
CritiQ: Mining Data Quality Criteria from Human Preferences · ACL (1) 2025
Natural language and speech › Information extraction and text analysis
named entity recognition
0.712023
UTC-IE: A Unified Token-pair Classification Architecture for Information Extraction · ACL (1) 2023
Natural language and speech › Information extraction and text analysis
relation extraction
0.712023
UTC-IE: A Unified Token-pair Classification Architecture for Information Extraction · ACL (1) 2023
Natural language and speech › Information extraction and text analysis
span extraction
0.712023
UTC-IE: A Unified Token-pair Classification Architecture for Information Extraction · ACL (1) 2023
Natural language and speech › Information extraction and text analysis
event extraction
0.212023
UTC-IE: A Unified Token-pair Classification Architecture for Information Extraction · ACL (1) 2023

Methods — techniques the papers use, named apart from their topics

preference mining · 1.7plus-shaped self-attention · 0.7convolutional neural network · 0.7
YearPublicationVenuePosition
2025 CritiQ: Mining Data Quality Criteria from Human Preferences
abstract
Honglin Guo, Kai Lv, Qipeng Guo, Tianyi Liang, Zhiheng Xi, Demin Song, Qiuyinzhe Zhang, Yu Sun, Kai Chen, Xipeng Qiu, Tao Gui. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Honglin Guo, Kai Lv 0001, Qipeng Guo, Tianyi Liang 0002, Zhiheng Xi, Demin Song, Qiuyinzhe Zhang, Yu Sun 0031, Kai Chen 0026, Xipeng Qiu, Tao Gui
ACL (1)8
2024 F-Eval: Asssessing Fundamental Abilities with Refined Evaluation Methods
abstract
Yu Sun, Keyu Chen, Shujie Wang, Peiji Li, Qipeng Guo, Hang Yan, Xipeng Qiu, Xuanjing Huang, Dahua Lin. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Yu Sun 0031, Keyuchen Keyuchen, Peiji Li, Qipeng Guo, Hang Yan 0001, Xipeng Qiu, Xuanjing Huang 0001, Dahua Lin
ACL (1)1
2023 UTC-IE: A Unified Token-pair Classification Architecture for Information Extraction
abstract
Information Extraction (IE) spans several tasks with different output structures, such as named entity recognition, relation extraction and event extraction.Previously, those tasks were solved with different models because of diverse task output structures.Through re-examining IE tasks, we find that all of them can be interpreted as extracting spans and span relations.They can further be decomposed into tokenpair classification tasks by using the start and end token of a span to pinpoint the span, and using the start-to-start and end-to-end token pairs of two spans to determine the relation.Based on the reformulation, we propose a Unified Token-pair Classification architecture for Information Extraction (UTC-IE), where we introduce Plusformer on top of the tokenpair feature matrix.Specifically, it models axis-aware interaction with plus-shaped selfattention and local interaction with Convolutional Neural Network over token pairs.Experiments show that our approach outperforms task-specific and unified models on all tasks in 10 datasets, and achieves better or comparable results on 2 joint IE datasets.Moreover, UTC-IE speeds up over state-of-the-art models on IE tasks significantly in most datasets, which verifies the effectiveness of our architecture.1 * Equal contribution.
Hang Yan 0001, Yu Sun 0031, Yunhua Zhou, Xuanjing Huang 0001, Xipeng Qiu
ACL (1)2
2022 BART-Reader: Predicting Relations Between Entities via Reading Their Document-Level Context Information
Hang Yan 0001, Yu Sun 0031, Junqi Dai, Xiangkun Hu, Qipeng Guo, Xipeng Qiu, Xuanjing Huang 0001
NLPCC (1)2