Huiyao Chen

dblp:359/7464 · DBLP profile ↗
← Back
5ranked-venue papers
4as first author
5since 2021 · last 2026
0000-0003-3518-0817ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 4 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Information extraction and text analysis · 67% Trustworthy machine learning · 33%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval
question answering
1.012026
Beyond Chunking: Discourse-Aware Hierarchical Retrieval for Long Document Question Answering · ACL (1) 2026
Machine learning › Trustworthy machine learning › robustness
learning with noisy labels
0.812024
Automatic Noise Generation and Reduction for Text Classification · IEEE ACM Trans. Audio Speech Lang. Process. 2024
Natural language and speech › Information extraction and text analysis › text classification › robust text classification
noisy text classification
0.812024
Automatic Noise Generation and Reduction for Text Classification · IEEE ACM Trans. Audio Speech Lang. Process. 2024
Natural language and speech › Information extraction and text analysis
text classification
0.812024
Automatic Noise Generation and Reduction for Text Classification · IEEE ACM Trans. Audio Speech Lang. Process. 2024
Information retrieval › document retrieval › structure-aware retrieval
hierarchical retrieval
0.312026
Beyond Chunking: Discourse-Aware Hierarchical Retrieval for Long Document Question Answering · ACL (1) 2026

Methods — techniques the papers use, named apart from their topics

rhetorical structure theory · 1.0large language model · 1.0discourse parsing · 1.0pseudo noise generation · 0.8noise reduction · 0.8crowdsourcing noise · 0.8
YearPublicationVenuePosition
2026 Beyond Chunking: Discourse-Aware Hierarchical Retrieval for Long Document Question Answering
abstract
Existing long-document question answering systems typically process texts as flat sequences or use heuristic chunking, which overlook the discourse structures that naturally guide human comprehension.We present a discourse-aware hierarchical framework that leverages rhetorical structure theory (RST) for long document question answering.Our approach converts discourse trees into sentence-level representations and employs LLM-enhanced node representations to bridge structural and semantic information.The framework involves three key innovations: language-universal discourse parsing for lengthy documents, LLM-based enhancement of discourse relation nodes, and structure-guided hierarchical retrieval.Extensive experiments on four datasets demonstrate consistent improvements over existing approaches through the incorporation of discourse structure, across multiple genres and languages.Moreover, the proposed framework exhibits strong robustness across diverse document types and linguistic settings.
Huiyao Chen, Meishan Zhang, Baotian Hu, Min Zhang 0005
ACL (1)1
2025 Dependency Scoring Learning and Corpus Boosting for Translation-Based Cross-Lingual Dependency Parsing
abstract
Dependency parsing is a fundamental task in natural language processing that involves identifying the grammatical relationships between words in a sentence. One promising approach for performing this task in languages lacking annotated treebanks is treebank translation, which utilizes word alignments to map dependencies from a source treebank to the corresponding target translation. However, due to language differences and the limitations of word alignment tools, this method would inevitably generate noise during mapping. To reduce the effect of noise, we first exploit MetaNet to compute quality scores for each dependency and identify low-score ones as noise. MetaNet is a fake teacher that learns to score homework (dependencies) by comparing answers from the top student (strong parser) and the regular student (weak parser) without knowing the correct answer (gold-standard). With the scoring capability of MetaNet, we design an iterative algorithm to boost the target treebank quality, which trains with high-quality dependencies and relabels the low-quality dependencies. Our method achieves better results than the originally translated treebanks and shows highly competitive performances with prior methods on the Universal Dependency Treebanks v2.2. We also provide detailed analysis and discussions.
Huiyao Chen, Xin Zhang 0097, Jing Chen 0062, Meishan Zhang, Min Zhang 0005
ACM Trans. Asian Low Resour. Lang. Inf. Process.1
2024 LLM-Driven Multimodal Opinion Expression Identification
Bonian Jia, Huiyao Chen, Yueheng Sun, Meishan Zhang, Min Zhang 0005
INTERSPEECH2
2024 Retrieval-style In-context Learning for Few-shot Hierarchical Text Classification
abstract
Abstract Hierarchical text classification (HTC) is an important task with broad applications, and few-shot HTC has gained increasing interest recently. While in-context learning (ICL) with large language models (LLMs) has achieved significant success in few-shot learning, it is not as effective for HTC because of the expansive hierarchical label sets and extremely ambiguous labels. In this work, we introduce the first ICL-based framework with LLM for few-shot HTC. We exploit a retrieval database to identify relevant demonstrations, and an iterative policy to manage multi-layer hierarchical labels. Particularly, we equip the retrieval database with HTC label-aware representations for the input texts, which is achieved by continual training on a pretrained language model with masked language modeling (MLM), layer-wise classification (CLS, specifically for HTC), and a novel divergent contrastive learning (DCL, mainly for adjacent semantically similar labels) objective. Experimental results on three benchmark datasets demonstrate superior performance of our method, and we can achieve state-of-the-art results in few-shot HTC.
Huiyao Chen, Yu Zhao 0043, Zulong Chen, Mengjia Wang, Liangyue Li, Meishan Zhang, Min Zhang 0005
Trans. Assoc. Comput. Linguistics1
2024 Automatic Noise Generation and Reduction for Text Classification
abstract
Label noise is an important issue in machine learning, which might lead to negative influences on various tasks. Given that real benchmarks for evaluation of noise reduction methods are limited, plenty of studies construct pseudo noisy data to verify their proposed methods. However, very few works have realized the rationality of the noise generation strategies. If the generated pseudo datasets are biased, their final conclusions might also be problematic. In this work, we focus on text classification of natural language processing (NLP) to investigate various pseudo noise generation methods, which is the first work of this line for NLP. In particular, we compare the noise generated with crowdsourcing noise, a kind of real noise as gold-standard, to evaluate these noise generation methods. After then, we measure and compare the performance of representative noise reduction methods respectively based on the data of crowdsourcing and our top-ranked pseudo noisy generation strategies. We conduct experiments on five text classification datasets, offering detailed comparison results as well as discussions.
Huiyao Chen, Yueheng Sun, Meishan Zhang, Min Zhang 0005
IEEE ACM Trans. Audio Speech Lang. Process.1