Bairu Hou

dblp:274/7151 · DBLP profile ↗
← Back
7ranked-venue papers
6as first author
6since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 6 first-author · 6 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Efficient and distributed learning · 34% Trustworthy machine learning · 27% Language models and text generation · 26%

Topics — the 19 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
inference efficiency
0.912025
KVLink: Accelerating Large Language Models via Efficient KV Cache Reuse · NeurIPS 2025
Machine learning › Efficient and distributed learning
KV cache
0.912025
KVLink: Accelerating Large Language Models via Efficient KV Cache Reuse · NeurIPS 2025
Machine learning › Efficient and distributed learning › inference efficiency
KV cache reuse
0.912025
KVLink: Accelerating Large Language Models via Efficient KV Cache Reuse · NeurIPS 2025
Natural language and speech › Language models and text generation › efficient language model
large language model efficiency
0.912025
Instruction-Following Pruning for Large Language Models · ICML 2025
Natural language and speech › Language models and text generation › large language model inference
long-context inference
0.912025
KVLink: Accelerating Large Language Models via Efficient KV Cache Reuse · NeurIPS 2025
Machine learning › Efficient and distributed learning
model compression
0.912025
Instruction-Following Pruning for Large Language Models · ICML 2025
Machine learning › Efficient and distributed learning › model compression › pruning
structured pruning
0.912025
Instruction-Following Pruning for Large Language Models · ICML 2025
Machine learning › Trustworthy machine learning › uncertainty estimation
aleatoric and epistemic uncertainty
0.812024
Decomposing Uncertainty for Large Language Models through Input Clarification Ensembling · ICML 2024
Natural language and speech › Language models and text generation › trustworthy language model
large language model reliability
0.812024
Decomposing Uncertainty for Large Language Models through Input Clarification Ensembling · ICML 2024
Machine learning › Trustworthy machine learning › uncertainty estimation
uncertainty decomposition
0.812024
Decomposing Uncertainty for Large Language Models through Input Clarification Ensembling · ICML 2024
Machine learning › Trustworthy machine learning
uncertainty estimation
0.812024
Decomposing Uncertainty for Large Language Models through Input Clarification Ensembling · ICML 2024
Machine learning › Transfer learning and domain adaptation
few-shot learning
0.712023
PromptBoosting: Black-Box Text Classification with Ten Forward Passes · ICML 2023
Natural language and speech › Language models and text generation › trustworthy language model
natural language processing robustness
0.712023
TextGrad: Advancing Robustness Evaluation in NLP by Gradient-Driven Optimization · ICLR 2023
Computer vision › Vision and language › vision-language model
prompt learning
0.712023
PromptBoosting: Black-Box Text Classification with Ten Forward Passes · ICML 2023
Machine learning › Trustworthy machine learning
robustness
0.712023
TextGrad: Advancing Robustness Evaluation in NLP by Gradient-Driven Optimization · ICLR 2023
Machine learning › Trustworthy machine learning
robustness evaluation
0.712023
TextGrad: Advancing Robustness Evaluation in NLP by Gradient-Driven Optimization · ICLR 2023
Machine learning › Efficient and distributed learning › model compression › pruning
dynamic pruning
0.312025
Instruction-Following Pruning for Large Language Models · ICML 2025
Natural language and speech › Language models and text generation
retrieval-augmented generation
0.312025
KVLink: Accelerating Large Language Models via Efficient KV Cache Reuse · NeurIPS 2025
Natural language and speech › Information extraction and text analysis
text classification
0.212023
PromptBoosting: Black-Box Text Classification with Ten Forward Passes · ICML 2023

Methods — techniques the papers use, named apart from their topics

trainable special tokens · 0.9sparse mask prediction · 0.9positional embedding adjustment · 0.9joint optimization · 0.9input clarification ensembling · 0.8gradient-free prompt search · 0.7gradient-driven optimization · 0.7adaboost · 0.7
YearPublicationVenuePosition
2025 Instruction-Following Pruning for Large Language Models
abstract
With the rapid scaling of large language models (LLMs), structured pruning has become a widely used technique to learn efficient, smaller models from larger ones, delivering superior performance compared to training similarly sized models from scratch. In this paper, we move beyond the traditional static pruning approach of determining a fixed pruning mask for a model, and propose a dynamic approach to structured pruning. In our method, the pruning mask is input-dependent and adapts dynamically based on the information described in a user instruction. Our approach, termed "instruction-following pruning'', introduces a sparse mask predictor that takes the user instruction as input and dynamically selects the most relevant model parameters for the given task. To identify and activate effective parameters, we jointly optimize the sparse mask predictor and the LLM, leveraging both instruction-following data and the pre-training corpus. Experimental results demonstrate the effectiveness of our approach on a wide range of evaluation benchmarks. For example, our 3B activated model improves over the 3B dense model by 5-8 points of absolute margin on domains such as math and coding, and rivals the performance of a 9B model.
Bairu Hou, Guoli Yin, Nan Du 0002, Ruoming Pang, Shiyu Chang, Tao Lei 0001
ICML1
2025 A Probabilistic Framework for LLM Hallucination Detection via Belief Tree Propagation
abstract
Bairu Hou, Yang Zhang, Jacob Andreas, Shiyu Chang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Bairu Hou, Yang Zhang 0001, Jacob Andreas, Shiyu Chang
NAACL (Long Papers)1
2025 KVLink: Accelerating Large Language Models via Efficient KV Cache Reuse
abstract
We describe KVLink, an approach for efficient key-value (KV) cache reuse in large language models (LLMs). In many LLM applications, different inputs can share overlapping context, such as the same retrieved document appearing in multiple queries. However, the LLMs still need to encode the entire context for each query, leading to redundant computation. In this paper, we investigate a new strategy to eliminate such inefficiency, where the KV cache of each document is precomputed independently. During inference, the KV caches of retrieved documents are concatenated, allowing the model to reuse cached representations instead of recomputing them. To mitigate the performance degradation when using KV caches computed independently for each document, KVLink introduces two key techniques: adjusting positional embeddings of the KV cache at inference to match the global position after concatenation, and using trainable special tokens to restore self-attention across independently encoded documents. Experiments across 7 datasets demonstrate that KVLink improves question answering accuracy by an average of 4% over state-of-the-art methods. Furthermore, by leveraging precomputed KV caches, our approach reduces time-to-first-token by up to 96% compared to standard LLM inference, making it a scalable and efficient solution for context reuse. Additionally, KVLink can be combined with KV cache compression to further save cache loading and storage overhead while outperforming the baselines.
Bairu Hou, Wei Wei 0019, Yujia Bao, Shiyu Chang
NeurIPS2
2024 Decomposing Uncertainty for Large Language Models through Input Clarification Ensembling
abstract
Uncertainty decomposition refers to the task of decomposing the total uncertainty of a predictive model into aleatoric (data) uncertainty, resulting from inherent randomness in the data-generating process, and epistemic (model) uncertainty, resulting from missing information in the model's training data. In large language models (LLMs) specifically, identifying sources of uncertainty is an important step toward improving reliability, trustworthiness, and interpretability, but remains an important open research question. In this paper, we introduce an uncertainty decomposition framework for LLMs, called input clarification ensembling, which can be applied to any pre-trained LLM. Our approach generates a set of clarifications for the input, feeds them into an LLM, and ensembles the corresponding predictions. We show that, when aleatoric uncertainty arises from ambiguity or under-specification in LLM inputs, this approach makes it possible to factor an (un-clarified) LLM's predictions into separate aleatoric and epistemic terms, using a decomposition similar to the one employed by Bayesian neural networks. Empirical evaluations demonstrate that input clarification ensembling provides accurate and reliable uncertainty quantification on several language processing tasks. Code and data are available at https://github.com/UCSB-NLP-Chang/llm_uncertainty.
Bairu Hou, Yujian Liu, Kaizhi Qian, Jacob Andreas, Shiyu Chang, Yang Zhang 0001
ICML1
2023 TextGrad: Advancing Robustness Evaluation in NLP by Gradient-Driven Optimization
Bairu Hou, Jinghan Jia, Yang Zhang 0001, Sijia Liu 0001, Shiyu Chang
ICLR1
2023 PromptBoosting: Black-Box Text Classification with Ten Forward Passes
abstract
We describe PromptBoosting, a query-efficient procedure for building a text classifier from a neural language model (LM) without access to the LM’s parameters, gradients, or hidden representations. This form of "black-box" classifier training has become increasingly important as the cost of training and inference in large-scale LMs has grown. But existing black-box LM classifier learning approaches are themselves computationally inefficient, typically specializing LMs to the target task by searching in a large space of (discrete or continuous) prompts using zeroth-order optimization methods. Instead of directly optimizing in prompt space, PromptBoosting obtains a small pool of prompts via a gradient-free approach and then constructs a large pool of weak learners by pairing these prompts with different elements of the LM’s output distribution. These weak learners are then ensembled using the AdaBoost algorithm. The entire learning process requires only a small number of forward passes and no backward pass. Experiments show that PromptBoosting achieves state-of-the-art performance in multiple black-box few-shot classification tasks, and matches or outperforms full fine-tuning in both few-shot and standard learning paradigms, while training 10x faster than existing black-box methods.
Bairu Hou, Joe O'Connor, Jacob Andreas, Shiyu Chang, Yang Zhang 0001
ICML1
2020 Try to Substitute: An Unsupervised Chinese Word Sense Disambiguation Method Based on HowNet
abstract
Word sense disambiguation (WSD) is a fundamental natural language processing task.Unsupervised knowledge-based WSD only relies on a lexical knowledge base as the sense inventory and has wider practical use than supervised WSD that requires a mass of sense-annotated data.HowNet is the most widely used lexical knowledge base in Chinese WSD.Because of its uniqueness, however, most of existing unsupervised WSD methods cannot work for HowNetbased WSD, and the tailor-made methods have not obtained satisfying results.In this paper, we propose a new unsupervised method for HowNet-based Chinese WSD, which exploits the masked language model task of pre-trained language models.In experiments, considering existing evaluation dataset is small and out-of-date, we build a new and larger HowNet-based WSD dataset.Experimental results demonstrate that our model achieves significantly better performance than all the baseline methods.All the code and data of this paper are available at https://github.com/thunlp/SememeWSD.
Bairu Hou, Fanchao Qi, Yuan Zang, Xurui Zhang, Zhiyuan Liu 0001, Maosong Sun 0001
COLING1