Qianle Wang

dblp:358/6959 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Efficient and distributed learning · 76% Language models and text generation · 19% Deep learning architectures and training · 6%

Topics — the 5 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › efficient language model
large language model efficiency
0.812024
Cherry on Top: Parameter Heterogeneity and Quantization in Large Language Models · NeurIPS 2024
Machine learning › Efficient and distributed learning › model compression › quantization
mixed-precision quantization
0.812024
Cherry on Top: Parameter Heterogeneity and Quantization in Large Language Models · NeurIPS 2024
Machine learning › Efficient and distributed learning
model compression
0.812024
Cherry on Top: Parameter Heterogeneity and Quantization in Large Language Models · NeurIPS 2024
Machine learning › Efficient and distributed learning › model compression › quantization
post-training quantization
0.812024
Cherry on Top: Parameter Heterogeneity and Quantization in Large Language Models · NeurIPS 2024
Machine learning › Efficient and distributed learning › model compression
quantization
0.812024
Cherry on Top: Parameter Heterogeneity and Quantization in Large Language Models · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

parameter importance analysis · 0.8mixed-precision quantization · 0.8
YearPublicationVenuePosition
2024 Cherry on Top: Parameter Heterogeneity and Quantization in Large Language Models
abstract
This paper reveals the phenomenon of parameter heterogeneity in large language models (LLMs). We find that a small subset of ``cherry'' parameters exhibit a disproportionately large influence on model performance, while the vast majority of parameters have minimal impact. This heterogeneity is found to be prevalent across different model families, scales, and types. Motivated by this observation, we propose CherryQ, a novel quantization method that unifies the optimization of mixed-precision parameters. CherryQ identifies and preserves the critical cherry parameters in high precision while aggressively quantizing the remaining parameters to low precision. Extensive experiments demonstrate the effectiveness of CherryQ. CherryQ outperforms existing quantization approaches in terms of perplexity and downstream task performance. Notably, our 3-bit quantized Vicuna-1.5 exhibits competitive performance compared to their 16-bit counterparts. These findings highlight the potential of CherryQ for enabling efficient deployment of LLMs by taking advantage of parameter heterogeneity.
Wanyun Cui, Qianle Wang
NeurIPS2