Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Yueqi Zhou 0001

dblp:166/7002-1 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2026
0009-0009-7531-6396ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Efficient and distributed learning · 100%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning › model compression
large language model compression
1.012026
Sliding-Window Merging for Compacting Patch-Redundant Layers in LLMs · AAAI 2026
Machine learning › Efficient and distributed learning › model merging
layer merging
1.012026
Sliding-Window Merging for Compacting Patch-Redundant Layers in LLMs · AAAI 2026
Machine learning › Efficient and distributed learning › model compression › pruning › structured pruning
layer pruning
1.012026
Sliding-Window Merging for Compacting Patch-Redundant Layers in LLMs · AAAI 2026
Machine learning › Efficient and distributed learning
model compression
1.012026
Sliding-Window Merging for Compacting Patch-Redundant Layers in LLMs · AAAI 2026

Methods — techniques the papers use, named apart from their topics

sliding-window merging · 1.0reproducing kernel hilbert space · 1.0correlation analysis · 1.0
YearPublicationVenuePosition
2026 Sliding-Window Merging for Compacting Patch-Redundant Layers in LLMs
abstract
Depth-wise pruning accelerates LLM inference in resource-constrained scenarios but suffers from performance degradation due to indiscriminate removal of entire Transformer layers. This paper reveals ``Patch-Like'' redundancy across layers via correlation analysis of the outputs of different layers in reproducing kernel Hilbert space, demonstrating consecutive layers exhibit high functional similarity. Building on this observation, this paper proposes Sliding-Window Merging (SWM) - a dynamic compression method that selects consecutive layers from top to bottom using a pre-defined similarity threshold, and compacts patch-redundant layers through a parameter consolidation, thereby simplifying the model structure while maintaining its performance. Extensive experiments on LLMs with various architectures and different parameter scales show that our method outperforms existing pruning techniques in both zero-shot inference performance and retraining recovery quality after pruning. In particular, in the experiment with 35\% pruning on the Vicuna-7B model, our method achieved a 1.654\% improvement in average performance on zero-shot tasks compared to the existing method. Moreover, we further reveal the potential of combining depth pruning with width pruning to enhance the pruning effect.
Xiu Yan, Yueqi Zhou 0001, Kaihao Huang, Suzhong Fu, Angelica I. Avilés-Rivero, Chuanlong Xie, Yao Zhu 0003
AAAI5