Bo Wang 0084

dblp:72/6811-84 · DBLP profile ↗
← Back
7ranked-venue papers
1as first author
7since 2021 · last 2026
0000-0003-0526-0533ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Security and privacy · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Language models and text generation · 64% Efficient and distributed learning · 18% Reinforcement learning · 16%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 20 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
large language model
1.222026
BitStack: Any-Size Compression of Large Language Models in Variable Memory Environments · ICLR 2025
Efficient KL Divergence Estimation via Truncated Top-K Integration for Large Language Models · ACL (1) 2026
Natural language and speech › Language models and text generation
chain-of-thought reasoning
1.012026
Time-Frequency Token Advantage Clipping for Training Efficient Large Reasoning Model · AAAI 2026
Natural language and speech › Language models and text generation › large language model reasoning
efficient reasoning
1.012026
Time-Frequency Token Advantage Clipping for Training Efficient Large Reasoning Model · AAAI 2026
Natural language and speech › Language models and text generation › large language model
large reasoning model
1.012026
Time-Frequency Token Advantage Clipping for Training Efficient Large Reasoning Model · AAAI 2026
Machine learning › Reinforcement learning
reinforcement learning from human feedback
1.012026
Efficient KL Divergence Estimation via Truncated Top-K Integration for Large Language Models · ACL (1) 2026
Natural language and speech › Language models and text generation
alignment
0.912025
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections · NeurIPS 2025
Natural language and speech › Language models and text generation › preference optimization
direct preference optimization
0.912025
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections · NeurIPS 2025
Machine learning › Efficient and distributed learning
model compression
0.912025
BitStack: Any-Size Compression of Large Language Models in Variable Memory Environments · ICLR 2025
Natural language and speech › Language models and text generation › large language model training
post-training
0.912025
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections · NeurIPS 2025
Machine learning › Reinforcement learning
preference learning
0.912025
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections · NeurIPS 2025
Natural language and speech › Language models and text generation › large language model › large language model adaptation
supervised fine-tuning
0.912025
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections · NeurIPS 2025
Machine learning › Efficient and distributed learning › model compression › quantization
weight quantization
0.912025
BitStack: Any-Size Compression of Large Language Models in Variable Memory Environments · ICLR 2025
Information retrieval › reranking
listwise reranking
0.912025
REARANK: Reasoning Re-ranking Agent via Reinforcement Learning · EMNLP 2025
Information retrieval
ranking
0.912025
REARANK: Reasoning Re-ranking Agent via Reinforcement Learning · EMNLP 2025
Information retrieval
reranking
0.912025
REARANK: Reasoning Re-ranking Agent via Reinforcement Learning · EMNLP 2025
Machine learning › Efficient and distributed learning
inference efficiency
0.812024
Memorize Step by Step: Efficient Long-Context Prefilling with Incremental Memory and Decremental Chunk · EMNLP 2024
Natural language and speech › Language models and text generation › large language model inference
long-context inference
0.812024
Memorize Step by Step: Efficient Long-Context Prefilling with Incremental Memory and Decremental Chunk · EMNLP 2024
Machine learning › Reinforcement learning
policy optimization
0.312026
Time-Frequency Token Advantage Clipping for Training Efficient Large Reasoning Model · AAAI 2026
Natural language and speech › Language models and text generation
large language model reasoning
0.312025
REARANK: Reasoning Re-ranking Agent via Reinforcement Learning · EMNLP 2025
Machine learning › Deep learning architectures and training
weight decomposition
0.312025
BitStack: Any-Size Compression of Large Language Models in Variable Memory Environments · ICLR 2025

Methods — techniques the papers use, named apart from their topics

reinforcement learning · 2.7data augmentation · 1.7top-k integration · 1.0importance sampling · 1.0advantage clipping · 1.0KL divergence estimation · 1.0weight decomposition · 0.9residual block stacking · 0.9f-divergence · 0.9KL divergence · 0.9
YearPublicationVenuePosition
2026 Time-Frequency Token Advantage Clipping for Training Efficient Large Reasoning Model
abstract
Long Chain-of-Thought (CoT) reasoning enhances large reasoning models' performance but suffers from severe inefficiencies, as models often overthink simple problems or underthink complex ones. Current sequence-level optimizations, like length penalties, are too coarse-grained to distinguish core logic from verbose language, precluding the necessary token-level control for efficient reasoning CoT. To overcome these limitations, we introduce Time-Frequency token Advantage Clipping (TFAC), a novel training framework designed to build efficient large reasoning models via token-level interventions. Specifically, TFAC functions along two dimensions: 1) The Frequency Dimension: It discourages inefficient loops and encourages deeper exploration by dynamically reducing the advantage scores of high-entropy tokens that are repeatedly generated within a single reasoning path. 2) The Time Dimension: It reduces excessive overthinking of the system by establishing a historical baseline for the occurrence count of each critical token in previously successful trajectories, and clipping the advantages of tokens that exceed this baseline during training. Crucially, to preserve the model's exploratory capabilities on novel problems, this suppression mechanism is automatically disabled when no historical record of success is available. Experiments conducted on the Deepseek-Distill-32B and Qwen3-8B models show that TFAC outperforms leading baseline methods, improving performance by 2.3 and 3.1 percentage points, respectively, while simultaneously reducing inference costs by 35% and 28% in scenarios where correct answers are generated. These results validate the significant efficacy of TFAC in training large reasoning models that are both powerful and highly efficient.
Rong Bao, Bo Wang 0084, Xiao Wang 0014, Hongyu Li 0004, Leszek Rutkowski, Qi Zhang 0001, Liang Ding 0006, Dacheng Tao
AAAI2
2026 Efficient KL Divergence Estimation via Truncated Top-K Integration for Large Language Models
abstract
Kullback-Leibler (KL) divergence regularization is essential for stabilizing reinforcement learning from human feedback (RLHF) in large language models (LLMs), yet its exact computation requires summing over vocabularies of all tokens, incurring prohibitive memory costs during training.Existing stochastic estimators circumvent this bottleneck by estimating KL divergence using only the sampled token from the trajectory, but suffer from high variance (k 1 ) or systematic bias (k 2 ).We propose TIKE (Top-k Importance-weighted KL Estimator), which exploits the Zipfian structure of language model distributions: by deterministically integrating over only the top-k tokens, TIKE captures most of the probability mass while effectively reducing memory cost.To ensure correctness in off-policy settings characteristic of Group Relative Policy Optimization (GRPO), we incorporate importance sampling weights that correct for distribution shift between rollout and optimization policies.Experiments on models across diverse benchmarks demonstrate that TIKE consistently outperforms stochastic baselines, while exhibiting substantially lower gradient variance.Our analysis reveals that TIKE closely tracks the exact Rao-Blackwellized estimator with nearzero variance, offering a practical path toward stable, memory-efficient KL regularization for reasoning-intensive LLMs training.Code:
Luozhijie Jin, Bo Wang 0084, Zhangyue Yin, Xipeng Qiu
ACL (1)3
2026 Verifiable and Fair Registered Attribute-Based Multi-Hop Proxy Re-Encryption Scheme for LLM Agents
abstract
As Large Language Model (LLM) agents emerge as intelligent coordinators, they intensify the demand for secure sharing and forwarding of sensitive data in data-driven systems. Existing Attribute-Based Proxy Re-Encryption (ABPRE) offers fine-grained access control and secure data forwarding, but suffers from the key escrow problem and lacks efficient verifiability and fairness guarantees. In this paper, we propose the first Verifiable and Fair Registered ABPRE (VF-RABPRE) scheme to eliminate reliance on trusted authorities and support multi-hop re-encryption for flexible multi-agent data sharing. To ensure efficient verifiability and fairness, we combine a lightweight verifiable tag and a non-interactive zero-knowledge proof to detect misbehavior of the proxy and prevent false accusations. As a trade-off between security and overhead, we design a more secure fairness mechanism that does not reveal plaintext using zero-knowledge Succinct Non-interactive Arguments of Knowledge (zkSNARK). Additionally, we extend VF-RABPRE with dynamic user registration and outsourced decryption, supporting flexible registration and efficient decryption for data users. Finally, we formally prove the security of our scheme and implement a prototype to evaluate its performance. Experimental results demonstrate that VF-RABPRE outperforms state-of-the-art ABPRE schemes and achieves practical efficiency, making it well-suited for secure data sharing in scenarios empowered by LLM agents.
Dongliang Cai, Yiwen Gao 0007, Qixiang Li, Liang Zhang 0043, Borui Chen, Bo Wang 0084, Haibin Kan
IEEE Trans. Inf. Forensics Secur.6
2025 REARANK: Reasoning Re-ranking Agent via Reinforcement Learning
abstract
We present REARANK, a large language model (LLM)-based listwise reasoning reranking agent.REARANK explicitly reasons before reranking, significantly improving both performance and interpretability.Leveraging reinforcement learning and data augmentation, REARANK achieves substantial improvements over baseline models across popular information retrieval benchmarks, notably requiring only 179 annotated samples.Built on top of Qwen2.5-7B,our REARANK-7B demonstrates performance comparable to GPT-4 on both indomain and out-of-domain benchmarks and even surpasses GPT-4 on reasoning-intensive BRIGHT benchmarks.These results underscore the effectiveness of our approach and highlight how reinforcement learning can enhance LLM reasoning capabilities in reranking.
Bo Wang 0084, Xipeng Qiu, Siva Reddy, Aishwarya Agrawal
EMNLP2
2025 BitStack: Any-Size Compression of Large Language Models in Variable Memory Environments
abstract
Large language models (LLMs) have revolutionized numerous applications, yet their deployment remains challenged by memory constraints on local devices. While scaling laws have enhanced LLM capabilities, the primary bottleneck has shifted from $\textit{capability}$ to $\textit{availability}$, emphasizing the need for efficient memory management. Traditional compression methods, such as quantization, often require predefined compression ratios and separate compression processes for each setting, complicating deployment in variable memory environments. In this paper, we introduce $\textbf{BitStack}$, a novel, training-free weight compression approach that enables megabyte-level trade-offs between memory usage and model performance. By leveraging weight decomposition, BitStack can dynamically adjust the model size with minimal transmission between running memory and storage devices. Our approach iteratively decomposes weight matrices while considering the significance of each parameter, resulting in an approximately 1-bit per parameter residual block in each decomposition iteration. These blocks are sorted and stacked in storage as basic transmission units, with different quantities loaded based on current memory availability. Extensive experiments across a wide range of tasks demonstrate that, despite offering fine-grained size control, BitStack consistently matches or surpasses strong quantization baselines, particularly at extreme compression ratios. To the best of our knowledge, this is the first decomposition-based method that effectively bridges the gap to practical compression techniques like quantization. Code is available at https://github.com/xinghaow99/BitStack.
Pengyu Wang 0006, Bo Wang 0084, Yunhua Zhou, Xipeng Qiu
ICLR3
2025 Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections
abstract
Post-training processes are essential phases in grounding pre-trained language models to real-world tasks, with learning from demonstrations or preference signals playing a crucial role in this adaptation. We present a unified theoretical framework bridging Supervised Fine-Tuning (SFT) and preference learning in Large Language Model (LLM) post-training. Through rigorous mathematical derivation, we demonstrate that both SFT and preference learning methods like Direct Preference Optimization (DPO) operate within the same optimal policy-reward subspace, with SFT representing a special case of implicit reward learning. Our analysis reveals a critical limitation in conventional SFT: the KL divergence term in distribution matching becomes constant with respect to the policy during optimization, failing to constrain model updates. To address this, we propose a simple yet effective learning rate reduction approach that yields significant performance improvements (up to \textbf{25\%} relative gain and \textbf{6\%} absolute win rate increase in instruction following tasks. Additionally, we derive alternative SFT objectives from various f-divergence functions that preserve the KL term during optimization, further enhancing post-DPO model performance. Finally, we extend the theoretical relationship between LLM logits and Q-functions from preference learning to the SFT context, providing mathematical derivations and experimental validation.
Bo Wang 0084, Qinyuan Cheng, Runyu Peng, Rong Bao, Peiji Li, Qipeng Guo, Linyang Li, Zhiyuan Zeng 0004, Yunhua Zhou, Xipeng Qiu
NeurIPS1
2024 Memorize Step by Step: Efficient Long-Context Prefilling with Incremental Memory and Decremental Chunk
abstract
Zhiyuan Zeng, Qipeng Guo, Xiaoran Liu, Zhangyue Yin, Wentao Shu, Mianqiu Huang, Bo Wang, Yunhua Zhou, Linlin Li, Qun Liu, Xipeng Qiu. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Zhiyuan Zeng 0004, Qipeng Guo, Zhangyue Yin, Wentao Shu, Mianqiu Huang, Bo Wang 0084, Yunhua Zhou, Linlin Li 0001, Qun Liu 0001, Xipeng Qiu
EMNLP7