VLDB 2026 Research / reviewers in the wild / expert
Baohao Liao
dblp:234/4096
· DBLP profile ↗
6ranked-venue papers
5as first author
5since 2021 · last 2025
0000-0001-8335-4573ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 5 first-author · 5 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Efficient and distributed learning · 65% Language models and text generation · 33% Trustworthy machine learning · 2% |
Topics — the 15 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning |
2.8 | 4 | 2024 | 3-in-1: 2D Rotary Adaptation for Efficient Finetuning, Efficient Batching and Composability · NeurIPS 2024 ApiQ: Finetuning of 2-Bit Quantized Large Language Model · EMNLP 2024 Make Pre-trained Model Reversible: From Parameter to Memory Efficient Fine-Tuning · NeurIPS 2023 |
Machine learning › Efficient and distributed learning
inference efficiency |
0.9 | 1 | 2025 | Reward-Guided Speculative Decoding for Efficient LLM Reasoning · ICML 2025 |
Natural language and speech › Language models and text generation
large language model inference |
0.9 | 1 | 2025 | Reward-Guided Speculative Decoding for Efficient LLM Reasoning · ICML 2025 |
Natural language and speech › Language models and text generation › decoding › decoding strategy
reward-guided decoding |
0.9 | 1 | 2025 | Reward-Guided Speculative Decoding for Efficient LLM Reasoning · ICML 2025 |
Machine learning › Efficient and distributed learning › inference acceleration
speculative decoding |
0.9 | 1 | 2025 | Reward-Guided Speculative Decoding for Efficient LLM Reasoning · ICML 2025 |
Machine learning › Efficient and distributed learning › parameter-efficient fine-tuning
adapter tuning |
0.8 | 1 | 2024 | 3-in-1: 2D Rotary Adaptation for Efficient Finetuning, Efficient Batching and Composability · NeurIPS 2024 |
Natural language and speech › Language models and text generation › large language model
large language model adaptation |
0.8 | 1 | 2024 | 3-in-1: 2D Rotary Adaptation for Efficient Finetuning, Efficient Batching and Composability · NeurIPS 2024 |
Natural language and speech › Language models and text generation
large language model fine-tuning |
0.8 | 1 | 2024 | ApiQ: Finetuning of 2-Bit Quantized Large Language Model · EMNLP 2024 |
Machine learning › Efficient and distributed learning
model compression |
0.8 | 1 | 2024 | ApiQ: Finetuning of 2-Bit Quantized Large Language Model · EMNLP 2024 |
Machine learning › Efficient and distributed learning › model compression
quantization |
0.8 | 1 | 2024 | ApiQ: Finetuning of 2-Bit Quantized Large Language Model · EMNLP 2024 |
Machine learning › Efficient and distributed learning › memory-efficient training
memory-efficient fine-tuning |
0.7 | 1 | 2023 | Make Pre-trained Model Reversible: From Parameter to Memory Efficient Fine-Tuning · NeurIPS 2023 |
Natural language and speech › Language models and text generation › large language model › large language model adaptation
pre-trained language model fine-tuning |
0.7 | 1 | 2023 | Make Pre-trained Model Reversible: From Parameter to Memory Efficient Fine-Tuning · NeurIPS 2023 |
Machine learning › Efficient and distributed learning › parameter-efficient fine-tuning
sparse fine-tuning |
0.7 | 1 | 2023 | Parameter-Efficient Fine-Tuning without Introducing New Latency · ACL (1) 2023 |
Machine learning › Trustworthy machine learning
interpretability |
0.2 | 1 | 2024 | 3-in-1: 2D Rotary Adaptation for Efficient Finetuning, Efficient Batching and Composability · NeurIPS 2024 |
Natural language and speech › Language models and text generation
pre-trained language model |
0.2 | 1 | 2023 | Parameter-Efficient Fine-Tuning without Introducing New Latency · ACL (1) 2023 |
Methods — techniques the papers use, named apart from their topics
adapter · 1.3threshold-based mixture strategy · 0.9process reward model · 0.9draft model · 0.9quantization-aware initialization · 0.8distributed interchange interventions · 0.8LoRA · 0.82d rotation adaptation · 0.8sparse mask generation · 0.7magnitude pruning · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Reward-Guided Speculative Decoding for Efficient LLM ReasoningabstractWe introduce Reward-Guided Speculative Decoding (RSD), a novel framework aimed at improving the efficiency of inference in large language models (LLMs). RSD synergistically combines a lightweight draft model with a more powerful target model, incorporating a controlled bias to prioritize high-reward outputs, in contrast to existing speculative decoding methods that enforce strict unbiasedness. RSD employs a process reward model to evaluate intermediate decoding steps and dynamically decide whether to invoke the target model, optimizing the trade-off between computational cost and output quality. We theoretically demonstrate that a threshold-based mixture strategy achieves an optimal balance between resource utilization and performance. Extensive evaluations on challenging reasoning benchmarks, including Olympiad-level tasks, show that RSD delivers significant efficiency gains against decoding with the target model only (up to 4.4X fewer FLOPs), while achieving significant better accuracy than parallel decoding method on average (up to +3.5). These results highlight RSD as a robust and cost-effective approach for deploying LLMs in resource-intensive scenarios. Baohao Liao, Hanze Dong, Junnan Li 0001, Christof Monz, Silvio Savarese, Doyen Sahoo, Caiming Xiong |
ICML | 1 |
| 2024 | ApiQ: Finetuning of 2-Bit Quantized Large Language ModelabstractMemory-efficient finetuning of large language models (LLMs) has recently attracted huge attention with the increasing size of LLMs, primarily due to the constraints posed by GPU memory limitations and the effectiveness of these methods compared to full finetuning.Despite the advancements, current strategies for memory-efficient finetuning, such as QLoRA, exhibit inconsistent performance across diverse bit-width quantizations and multifaceted tasks.This inconsistency largely stems from the detrimental impact of the quantization process on preserved knowledge, leading to catastrophic forgetting and undermining the utilization of pretrained models for finetuning purposes.In this work, we introduce a novel quantization framework named ApiQ, designed to restore the lost information from quantization by concurrently initializing the LoRA components and quantizing the weights of LLMs.This approach ensures the maintenance of the original LLM's activation precision while mitigating the error propagation from shallower into deeper layers.Through comprehensive evaluations conducted on a spectrum of language tasks with various LLMs, ApiQ demonstrably minimizes activation error during quantization.Consequently, it consistently achieves superior finetuning results across various bit-widths.Notably, one can even finetune a 2-bit Llama-2-70b with ApiQ on a single NVIDIA A100-80GB GPU without any memory-saving techniques, and achieve promising results. Baohao Liao, Christian Herold, Shahram Khadivi, Christof Monz |
EMNLP | 1 |
| 2024 | 3-in-1: 2D Rotary Adaptation for Efficient Finetuning, Efficient Batching and ComposabilityabstractParameter-efficient finetuning (PEFT) methods effectively adapt large language models (LLMs) to diverse downstream tasks, reducing storage and GPU memory demands. Despite these advantages, several applications pose new challenges to PEFT beyond mere parameter efficiency. One notable challenge involves the efficient deployment of LLMs equipped with multiple task- or user-specific adapters, particularly when different adapters are needed for distinct requests within the same batch. Another challenge is the interpretability of LLMs, which is crucial for understanding how LLMs function. Previous studies introduced various approaches to address different challenges. In this paper, we introduce a novel method, RoAd, which employs a straightforward 2D rotation to adapt LLMs and addresses all the above challenges: (1) RoAd is remarkably parameter-efficient, delivering optimal performance on GLUE, eight commonsense reasoning tasks and four arithmetic reasoning tasks with <0.1% trainable parameters; (2) RoAd facilitates the efficient serving of requests requiring different adapters within a batch, with an overhead comparable to element-wise multiplication instead of batch matrix multiplication; (3) RoAd enhances LLM's interpretability through integration within a framework of distributed interchange intervention, demonstrated via composition experiments. Baohao Liao, Christof Monz |
NeurIPS | 1 |
| 2023 | Parameter-Efficient Fine-Tuning without Introducing New LatencyabstractParameter-efficient fine-tuning (PEFT) of pretrained language models has recently demonstrated remarkable achievements, effectively matching the performance of full fine-tuning while utilizing significantly fewer trainable parameters, and consequently addressing the storage and communication constraints.Nonetheless, various PEFT methods are limited by their inherent characteristics.In the case of sparse fine-tuning, which involves modifying only a small subset of the existing parameters, the selection of fine-tuned parameters is task-and domain-specific, making it unsuitable for federated learning.On the other hand, PEFT methods with adding new parameters typically introduce additional inference latency.In this paper, we demonstrate the feasibility of generating a sparse mask in a task-agnostic manner, wherein all downstream tasks share a common mask.Our approach, which relies solely on the magnitude information of pre-trained parameters, surpasses existing methodologies by a significant margin when evaluated on the GLUE benchmark.Additionally, we introduce a novel adapter technique that directly applies the adapter to pre-trained parameters instead of the hidden representation, thereby achieving identical inference speed to that of full finetuning.Through extensive experiments, our proposed method attains a new state-of-the-art outcome in terms of both performance and storage efficiency, storing only 0.03% parameters of full fine-tuning.1 Baohao Liao, Christof Monz |
ACL (1) | 1 |
| 2023 | Make Pre-trained Model Reversible: From Parameter to Memory Efficient Fine-TuningabstractParameter-efficient fine-tuning (PEFT) of pre-trained language models (PLMs) has emerged as a highly successful approach, with training only a small number of parameters without sacrificing performance and becoming the de-facto learning paradigm with the increasing size of PLMs. However, existing PEFT methods are not memory-efficient, because they still require caching most of the intermediate activations for the gradient calculation, akin to fine-tuning. One effective way to reduce the activation memory is to apply a reversible model, so the intermediate activations are not necessary to be cached and can be recomputed. Nevertheless, modifying a PLM to its reversible variant is not straightforward, since the reversible model has a distinct architecture from the currently released PLMs. In this paper, we first investigate what is a key factor for the success of existing PEFT methods, and realize that it's essential to preserve the PLM's starting point when initializing a PEFT method. With this finding, we propose memory-efficient fine-tuning (MEFT) that inserts adapters into a PLM, preserving the PLM's starting point and making it reversible without additional pre-training. We evaluate MEFT on the GLUE benchmark and five question-answering tasks with various backbones, BERT, RoBERTa, BART and OPT. MEFT significantly reduces the activation memory up to 84% of full fine-tuning with a negligible amount of trainable parameters. Moreover, MEFT achieves the same score on GLUE and a comparable score on the question-answering tasks as full fine-tuning. A similar finding is also observed for the image classification task. Baohao Liao, Shaomu Tan, Christof Monz |
NeurIPS | 1 |
| 2020 | Unifying Input and Output Smoothing in Neural Machine TranslationabstractSoft contextualized data augmentation is a recent method that replaces one-hot representation of words with soft posterior distributions of an external language model, smoothing the input of neural machine translation systems.Label smoothing is another effective method that penalizes over-confident model outputs by discounting some probability mass from the true target word, smoothing the output of neural machine translation systems.Having the benefit of updating all word vectors in each optimization step and better regularizing the models, the two smoothing methods are shown to bring significant improvements in translation performance.In this work, we study how to best combine the methods and stack the improvements.Specifically, we vary the prior distributions to smooth with, the hyperparameters that control the smoothing strength, and the token selection procedures.We conduct extensive experiments on small datasets, evaluate the recipes on larger datasets, and examine the implications when back-translation is further used.Our results confirm cumulative improvements when input and output smoothing are used in combination, giving up to +1.9 BLEU scores on standard machine translation tasks and reveal reasons why these smoothing methods should be preferred. Yingbo Gao, Baohao Liao, Hermann Ney |
COLING | 2 |