Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Kangyu Qiao

dblp:377/4953 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Efficient and distributed learning · 91% Language models and text generation · 9%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning › model compression › pruning › structured pruning
layer pruning
0.912025
IG-Pruning: Input-Guided Block Pruning for Large Language Models · EMNLP 2025
Machine learning › Efficient and distributed learning
model compression
0.912025
IG-Pruning: Input-Guided Block Pruning for Large Language Models · EMNLP 2025
Machine learning › Efficient and distributed learning › model compression › pruning
structured pruning
0.912025
IG-Pruning: Input-Guided Block Pruning for Large Language Models · EMNLP 2025
Natural language and speech › Language models and text generation
large language model inference
0.312025
IG-Pruning: Input-Guided Block Pruning for Large Language Models · EMNLP 2025

Methods — techniques the papers use, named apart from their topics

semantic clustering · 0.9l0 optimization · 0.9
YearPublicationVenuePosition
2025 IG-Pruning: Input-Guided Block Pruning for Large Language Models
abstract
With the growing computational demands of large language models (LLMs), efficient inference has become increasingly critical for practical deployment. Depth pruning has emerged as a promising approach for reducing the computational costs of large language models by removing transformer layers. However, existing methods typically rely on fixed block masks, which can lead to suboptimal performance across different tasks and inputs. In this paper, we propose IG-Pruning, a novel input-aware block-wise pruning method that dynamically selects layer masks at inference time. Our approach consists of two stages: (1) Discovering diverse mask candidates through semantic clustering and L0 optimization, and (2) Implementing efficient dynamic pruning without the need for extensive training. Experimental results demonstrate that our method consistently outperforms state-of-the-art static depth pruning methods, making it particularly suitable for resource-constrained deployment scenarios.
Kangyu Qiao, Shaolei Zhang 0001, Yang Feng 0004
EMNLP1
2024 Enhancing Document Information Selection Through Multi-Granularity Responses for Dialogue Generation
abstract
Abstract Document information selection is an essential part of document-grounded dialogue tasks, and more accurate information selection results can provide more appropriate dialogue responses. Existing works have achieved excellent results by employing multi-granularity of dialogue history information, indicating the effectiveness of multi-level historical information. However, these works often focus on exploring the hierarchical information of dialogue history, while neglecting the multi-granularity utilization in response, important information that holds an impact on the decoding process. Therefore, this paper proposes a model for document information selection based on multi-granularity responses. By integrating the document selection results at the response word level and semantic unit level, the model enhances its capability in knowledge selection and produces better responses. For the division at the semantic unit level of the response, we propose two semantic unit division methods, static and dynamic. Experiments on two public datasets show that our models combining static or dynamic semantic unit levels significantly outperform baseline models.
Kangyu Qiao, Shuyue Xing, Caixia Yuan, Xiaojie Wang 0006
Neural Process. Lett.2