Bowen Ping

dblp:338/6403 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Language models and text generation · 46% Efficient and distributed learning · 45% Transfer learning and domain adaptation · 9%

Topics — the 10 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
large language model
1.122026
Delta-CoMe: Training-Free Delta-Compression with Mixed-Precision for Large Language Models · NeurIPS 2024
DecIF: Improving Instruction-Following through Decomposition · ACL (1) 2026
Natural language and speech › Language models and text generation › instruction following
instruction decomposition
1.012026
DecIF: Improving Instruction-Following through Decomposition · ACL (1) 2026
Natural language and speech › Language models and text generation
instruction following
1.012026
DecIF: Improving Instruction-Following through Decomposition · ACL (1) 2026
Machine learning › Efficient and distributed learning › model compression
delta compression
0.812024
Delta-CoMe: Training-Free Delta-Compression with Mixed-Precision for Large Language Models · NeurIPS 2024
Machine learning › Transfer learning and domain adaptation
fine-tuning
0.812024
Delta-CoMe: Training-Free Delta-Compression with Mixed-Precision for Large Language Models · NeurIPS 2024
Natural language and speech › Language models and text generation › large language model
large language model adaptation
0.812024
LoRA-Flow: Dynamic LoRA Fusion for Large Language Models in Generative Tasks · ACL (1) 2024
Machine learning › Efficient and distributed learning › model merging
LoRA merging
0.812024
LoRA-Flow: Dynamic LoRA Fusion for Large Language Models in Generative Tasks · ACL (1) 2024
Machine learning › Efficient and distributed learning › model compression › quantization
mixed-precision quantization
0.812024
Delta-CoMe: Training-Free Delta-Compression with Mixed-Precision for Large Language Models · NeurIPS 2024
Machine learning › Efficient and distributed learning
model compression
0.812024
Delta-CoMe: Training-Free Delta-Compression with Mixed-Precision for Large Language Models · NeurIPS 2024
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning
0.812024
LoRA-Flow: Dynamic LoRA Fusion for Large Language Models in Generative Tasks · ACL (1) 2024

Methods — techniques the papers use, named apart from their topics

decomposition · 1.0singular value decomposition · 0.8mixed-precision quantization · 0.8low-rank decomposition · 0.8dynamic weight fusion · 0.8LoRA · 0.8
YearPublicationVenuePosition
2026 DecIF: Improving Instruction-Following through Decomposition
abstract
Tingfeng Hui, Pengyu Zhu, Bowen Ping, Ling Tang, Guanting Dong, Yaqi Zhang, Sen Su. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Tingfeng Hui, Bowen Ping, Guanting Dong 0001, Sen Su
ACL (1)3
2024 LoRA-Flow: Dynamic LoRA Fusion for Large Language Models in Generative Tasks
abstract
LoRA employs lightweight modules to customize large language models (LLMs) for each downstream task or domain, where different learned additional modules represent diverse skills.Combining existing LoRA modules to address new tasks can enhance the reusability of learned LoRA modules, particularly beneficial for tasks with limited annotated data.Most prior works on LoRA combination primarily rely on task-level weights for each involved LoRA, making different examples and tokens share the same LoRA weights.However, in generative tasks, different tokens may necessitate diverse skills to manage.Taking the Chinese math task as an example, understanding the problem description may depend more on the Chinese LoRA, while the calculation part may rely more on the math LoRA.To this end, we propose LoRA-Flow, which utilizes dynamic weights to adjust the impact of different LoRA modules.The weights at each step are determined by a fusion gate with extremely few parameters, which can be learned with only 200 training examples.Experiments across six generative tasks demonstrate that our method consistently outperforms baselines with tasklevel fusion weights.This underscores the necessity of introducing dynamic fusion weights for LoRA combination. 1
Hanqing Wang 0003, Bowen Ping, Shuo Wang 0013, Xu Han 0007, Yun Chen 0007, Zhiyuan Liu 0001, Maosong Sun 0001
ACL (1)2
2024 Delta-CoMe: Training-Free Delta-Compression with Mixed-Precision for Large Language Models
abstract
Fine-tuning is a crucial process for adapting large language models (LLMs) to diverse applications. In certain scenarios, such as multi-tenant serving, deploying multiple LLMs becomes necessary to meet complex demands. Recent studies suggest decomposing a fine-tuned LLM into a base model and corresponding delta weights, which are then compressed using low-rank or low-bit approaches to reduce costs. In this work, we observe that existing low-rank and low-bit compression methods can significantly harm the model performance for task-specific fine-tuned LLMs (e.g., WizardMath for math problems). Motivated by the long-tail distribution of singular values in the delta weights, we propose a delta quantization approach using mixed-precision. This method employs higher-bit representation for singular vectors corresponding to larger singular values. We evaluate our approach on various fine-tuned LLMs, including math LLMs, code LLMs, chat LLMs, and even VLMs. Experimental results demonstrate that our approach performs comparably to full fine-tuned LLMs, surpassing both low-rank and low-bit baselines by a considerable margin. Additionally, we show that our method is compatible with various backbone LLMs, such as Llama-2, Llama-3, and Mistral, highlighting its generalizability.
Bowen Ping, Shuo Wang 0013, Hanqing Wang 0003, Xu Han 0007, Yuzhuang Xu, Yukun Yan, Yun Chen 0007, Baobao Chang, Zhiyuan Liu 0001, Maosong Sun 0001
NeurIPS1