Yibin Wang 0005

dblp:56/8805-5 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2025
0000-0001-6099-2628ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Efficient and distributed learning · 33% Trustworthy machine learning · 25% Reinforcement learning · 20%

Topics — the 11 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning › parameter-efficient fine-tuning
low-rank adaptation
1.622025
Training-Free Bayesianization for Low-Rank Adapters of Large Language Models · NeurIPS 2025
BLoB: Bayesian Low-Rank Adaptation by Backpropagation for Large Language Models · NeurIPS 2024
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning
1.622025
Training-Free Bayesianization for Low-Rank Adapters of Large Language Models · NeurIPS 2025
BLoB: Bayesian Low-Rank Adaptation by Backpropagation for Large Language Models · NeurIPS 2024
Machine learning › Trustworthy machine learning
uncertainty estimation
1.622025
Training-Free Bayesianization for Low-Rank Adapters of Large Language Models · NeurIPS 2025
BLoB: Bayesian Low-Rank Adaptation by Backpropagation for Large Language Models · NeurIPS 2024
Machine learning › Trustworthy machine learning › uncertainty estimation
bayesian uncertainty quantification
0.912025
Training-Free Bayesianization for Low-Rank Adapters of Large Language Models · NeurIPS 2025
Natural language and speech › Language models and text generation
large language model reasoning
0.912025
Improving Data Efficiency for LLM Reinforcement Fine-tuning Through Difficulty-targeted Online Data Selection and Rollout Replay · NeurIPS 2025
Machine learning › Reinforcement learning › reinforcement learning for NLP
reinforcement fine-tuning
0.912025
Improving Data Efficiency for LLM Reinforcement Fine-tuning Through Difficulty-targeted Online Data Selection and Rollout Replay · NeurIPS 2025
Machine learning › Reinforcement learning › sample efficiency
sample-efficient reinforcement learning
0.912025
Improving Data Efficiency for LLM Reinforcement Fine-tuning Through Difficulty-targeted Online Data Selection and Rollout Replay · NeurIPS 2025
Machine learning › Probabilistic and Bayesian machine learning › deep probabilistic models › bayesian deep learning
bayesian neural networks
0.812024
BLoB: Bayesian Low-Rank Adaptation by Backpropagation for Large Language Models · NeurIPS 2024
Machine learning › Reinforcement learning › off-policy reinforcement learning
experience replay
0.312025
Improving Data Efficiency for LLM Reinforcement Fine-tuning Through Difficulty-targeted Online Data Selection and Rollout Replay · NeurIPS 2025
Machine learning › Transfer learning and domain adaptation
fine-tuning
0.212024
BLoB: Bayesian Low-Rank Adaptation by Backpropagation for Large Language Models · NeurIPS 2024
Natural language and speech › Language models and text generation
large language model
0.212024
BLoB: Bayesian Low-Rank Adaptation by Backpropagation for Large Language Models · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

variational inference · 0.9rollout replay · 0.9bayesianization · 0.9attention-based difficulty estimation · 0.9KL regularization · 0.9low-rank adaptation · 0.8bayesian estimation · 0.8backpropagation · 0.8
YearPublicationVenuePosition
2025 Training-Free Bayesianization for Low-Rank Adapters of Large Language Models
abstract
Estimating the uncertainty of responses from Large Language Models (LLMs) remains a critical challenge. While recent Bayesian methods have demonstrated effectiveness in quantifying uncertainty through low-rank weight updates, they typically require complex fine-tuning or post-training procedures. In this paper, we propose **T**raining-**F**ree **B**ayesianization (**TFB**), a simple yet theoretically grounded framework that efficiently transforms trained low-rank adapters into Bayesian ones without additional training. TFB systematically searches for the maximally acceptable level of variance in the weight posterior, constrained within a family of low-rank isotropic Gaussian distributions. Our theoretical analysis shows that under mild conditions, this search process is equivalent to KL-regularized variational optimization, a generalized form of variational inference. Through comprehensive experiments, we show that TFB achieves superior uncertainty estimation and generalization compared to existing methods while eliminating the need for complex Bayesianization training procedures.
Haizhou Shi, Yibin Wang 0005, Ligong Han, Huan Zhang 0001, Hao Wang 0014
NeurIPS2
2025 Improving Data Efficiency for LLM Reinforcement Fine-tuning Through Difficulty-targeted Online Data Selection and Rollout Replay
abstract
Reinforcement learning (RL) has become an effective approach for fine-tuning large language models (LLMs), particularly to enhance their reasoning capabilities. However, RL fine-tuning remains highly resource-intensive, and existing work has largely overlooked the problem of data efficiency. In this paper, we propose two techniques to improve data efficiency in LLM RL fine-tuning: difficulty-targeted online data selection and rollout replay. We introduce the notion of adaptive difficulty to guide online data selection, prioritizing questions of moderate difficulty that are more likely to yield informative learning signals. To estimate adaptive difficulty efficiently, we develop an attention-based framework that requires rollouts for only a small reference set of questions. The adaptive difficulty of the remaining questions is then estimated based on their similarity to this set. To further reduce rollout cost, we introduce a rollout replay mechanism inspired by experience replay in traditional RL. This technique reuses recent rollouts, lowering per-step computation while maintaining stable updates. Experiments across 6 LLM-dataset combinations show that our method reduces RL fine-tuning time by 23% to 62% while reaching the same level of performance as the original GRPO algorithm. Our code repository is available at https://github.com/ASTRAL-Group/data-efficient-llm-rl/.
Yifan Sun 0010, Jingyan Shen, Yibin Wang 0005, Zhendong Wang 0005, Mingyuan Zhou, Huan Zhang 0001
NeurIPS3
2024 BLoB: Bayesian Low-Rank Adaptation by Backpropagation for Large Language Models
abstract
Large Language Models (LLMs) often suffer from overconfidence during inference, particularly when adapted to downstream domain-specific tasks with limited data. Previous work addresses this issue by employing approximate Bayesian estimation after the LLMs are trained, enabling them to quantify uncertainty. However, such post-training approaches' performance is severely limited by the parameters learned during training. In this paper, we go beyond post-training Bayesianization and propose Bayesian Low-Rank Adaptation by Backpropagation (BLoB), an algorithm that continuously and jointly adjusts both the mean and covariance of LLM parameters throughout the whole fine-tuning process. Our empirical results verify the effectiveness of BLoB in terms of generalization and uncertainty estimation, when evaluated on both in-distribution and out-of-distribution data.
Yibin Wang 0005, Haizhou Shi, Ligong Han, Dimitris N. Metaxas, Hao Wang 0014
NeurIPS1