VLDB 2026 Research / reviewers in the wild / expert
Haokun Liu
dblp:169/0460
· DBLP profile ↗
17ranked-venue papers
3as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
10 papers |
Language models and text generation · 33% Efficient and distributed learning · 16% Trustworthy machine learning · 14% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Performance modeling and evaluation · 100% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computational science and engineering · 100% | |
| Software engineering, system software, and programming languages
1 paper |
Requirements engineering and software design · 100% |
Topics — the 21 heaviest of 25, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning |
1.3 | 2 | 2024 | Learning to Route Among Specialized Experts for Zero-Shot Generalization · ICML 2024 Few-Shot Parameter-Efficient Fine-Tuning is Better and Cheaper than In-Context Learning · NeurIPS 2022 |
Natural language and speech › Language models and text generation
pre-trained language model |
0.9 | 3 | 2020 | Learning Which Features Matter: RoBERTa Acquires a Preference for Linguistic Generalizations (Eventually) · EMNLP (1) 2020 Investigating BERT's Knowledge of Language: Five Analysis Methods with NPIs · EMNLP/IJCNLP (1) 2019 Intermediate-Task Transfer Learning with Pretrained Language Models: When and Why Does It Work? · ACL 2020 |
Computer vision › 3D vision › geometric estimation › geometric model fitting
hypothesis generation |
0.9 | 1 | 2025 | Literature Meets Data: A Synergistic Approach to Hypothesis Generation · ACL (1) 2025 |
Machine learning › Trustworthy machine learning
interpretability |
0.9 | 1 | 2025 | Enhancing Training Data Attribution with Representational Optimization · NeurIPS 2025 |
Machine learning › Trustworthy machine learning › interpretability
training data attribution |
0.9 | 1 | 2025 | Enhancing Training Data Attribution with Representational Optimization · NeurIPS 2025 |
Computational science and engineering › AI for science
AI for scientific discovery |
0.9 | 1 | 2025 | Literature Meets Data: A Synergistic Approach to Hypothesis Generation · ACL (1) 2025 |
Natural language and speech › Language models and text generation
linguistic generalization |
0.8 | 2 | 2020 | Learning Which Features Matter: RoBERTa Acquires a Preference for Linguistic Generalizations (Eventually) · EMNLP (1) 2020 Investigating BERT's Knowledge of Language: Five Analysis Methods with NPIs · EMNLP/IJCNLP (1) 2019 |
Machine learning › Deep learning architectures and training › mixture of experts
expert routing |
0.8 | 1 | 2024 | Learning to Route Among Specialized Experts for Zero-Shot Generalization · ICML 2024 |
Machine learning › Transfer learning and domain adaptation
zero-shot transfer |
0.8 | 1 | 2024 | Learning to Route Among Specialized Experts for Zero-Shot Generalization · ICML 2024 |
Requirements engineering and software design › model-driven engineering
collaborative modeling |
0.7 | 1 | 2023 | Git-Theta: A Git Extension for Collaborative Development of Machine Learning Models · ICML 2023 |
Natural language and speech › Language models and text generation
in-context learning |
0.6 | 1 | 2022 | Few-Shot Parameter-Efficient Fine-Tuning is Better and Cheaper than In-Context Learning · NeurIPS 2022 |
Natural language and speech › Language models and text generation › large language model evaluation
NLP evaluation |
0.5 | 1 | 2021 | Comparing Test Sets with Item Response Theory · ACL/IJCNLP (1) 2021 |
Performance modeling and evaluation
benchmarking |
0.5 | 1 | 2021 | Comparing Test Sets with Item Response Theory · ACL/IJCNLP (1) 2021 |
Performance modeling and evaluation › statistical analysis
item response theory |
0.5 | 1 | 2021 | Comparing Test Sets with Item Response Theory · ACL/IJCNLP (1) 2021 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
commonsense reasoning |
0.4 | 1 | 2020 | Precise Task Formalization Matters in Winograd Schema Evaluations · EMNLP (1) 2020 |
Machine learning › Learning theory
inductive bias |
0.4 | 1 | 2020 | Learning Which Features Matter: RoBERTa Acquires a Preference for Linguistic Generalizations (Eventually) · EMNLP (1) 2020 |
Natural language and speech › Language models and text generation › pre-trained language model
RoBERTa |
0.4 | 1 | 2020 | Learning Which Features Matter: RoBERTa Acquires a Preference for Linguistic Generalizations (Eventually) · EMNLP (1) 2020 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › commonsense reasoning
winograd schema challenge |
0.4 | 1 | 2020 | Precise Task Formalization Matters in Winograd Schema Evaluations · EMNLP (1) 2020 |
Natural language and speech › Language models and text generation › pre-trained language model › knowledge probing
linguistic knowledge probing |
0.4 | 1 | 2019 | Investigating BERT's Knowledge of Language: Five Analysis Methods with NPIs · EMNLP/IJCNLP (1) 2019 |
Natural language and speech › Information extraction and text analysis › text classification
deception detection |
0.3 | 1 | 2025 | Literature Meets Data: A Synergistic Approach to Hypothesis Generation · ACL (1) 2025 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
self-supervised representation learning |
0.1 | 1 | 2020 | Learning Which Features Matter: RoBERTa Acquires a Preference for Linguistic Generalizations (Eventually) · EMNLP (1) 2020 |
Methods — techniques the papers use, named apart from their topics
literature-based retrieval · 1.7data-driven generation · 1.7version control · 1.3model merging · 1.3ranking objective · 0.9influence functions · 0.9fine-tuning · 0.9attention-based pooling · 0.9tokenwise gating · 0.8parameter-efficient fine-tuning · 0.8item response theory · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Hierarchical Trajectory Planning of Floating-Base Multi-Link Robot for Maneuvering in Confined EnvironmentsabstractFloating-base multi-link robots can change their shape during flight, making them well-suited for applications in confined environments such as autonomous inspection and search and rescue. However, trajectory planning for such systems remains an open challenge because the problem lies in a high-dimensional, constraint-rich space where collision avoidance must be addressed together with kinematic limits and dynamic feasibility. This work introduces a hierarchical trajectory planning framework that integrates global guidance with configuration-aware local optimization. First, we exploit the dual nature of these robots—the root link as a rigid body for guidance and the articulated joints for flexibility—to generate global anchor states that decompose the planning problem into tractable segments. Second, we design a local trajectory planner that optimizes each segment in parallel with differentiable objectives and constraints, systematically enforcing kinematic feasibility and maintaining dynamic feasibility by avoiding control singularities. Third, we implement a complete system that directly processes point-cloud data, eliminating the need for handcrafted obstacle models. Extensive simulations and real-world experiments confirm that this framework enables an articulated aerial robot to exploit its morphology for maneuvering that rigid robots cannot achieve. To the best of our knowledge, this is the first planning framework for floating-base multi-link robots that has been demonstrated on a real robot to generate continuous, collision-free, and dynamically feasible trajectories directly from raw point-cloud inputs, without relying on handcrafted obstacle models. Jinjie Li, Haokun Liu, Zicheng Luo, Kotaro Kaneko, Moju Zhao |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2025 | Literature Meets Data: A Synergistic Approach to Hypothesis GenerationabstractAI holds promise for transforming scientific processes, including hypothesis generation.Prior work on hypothesis generation can be broadly categorized into theory-driven and datadriven approaches.While both have proven effective in generating novel and plausible hypotheses, it remains an open question whether they can complement each other.To address this, we develop the first method that combines literature-based insights with data to perform LLM-powered hypothesis generation.We apply our method on five different datasets and demonstrate that integrating literature and data outperforms other baselines (8.97% over fewshot, 15.75% over literature-based alone, and 3.37% over data-driven alone).Additionally, we conduct the first human evaluation to assess the utility of LLM-generated hypotheses in assisting human decision-making on two challenging tasks: deception detection and AI generated content detection.Our results show that human accuracy improves significantly by 7.44% and 14.19% on these tasks, respectively.These findings suggest that integrating literature-based and data-driven approaches provides a comprehensive and nuanced framework for hypothesis generation and could open new avenues for scientific inquiry. Haokun Liu, Yangqiaoyu Zhou, Chenfei Yuan, Chenhao Tan |
ACL (1) | 1 |
| 2025 | Six-DoF Hand-Based Teleoperation for Omnidirectional Aerial RobotsabstractOmnidirectional aerial robots offer full 6-DoF independent control over position and orientation, making them popular for aerial manipulation. Although advancements in robotic autonomy, human operation remains essential in complex aerial environments. Existing teleoperation approaches for multirotors fail to fully leverage the additional DoFs provided by omnidirectional rotation. Additionally, the dexterity of human fingers should be exploited for more engaged interaction. In this work, we propose an aerial teleoperation system that brings the rotational flexibility of human hands into the unbounded aerial workspace. Our system includes two motion-tracking marker sets—one on the shoulder and one on the hand—along with a data glove to capture hand gestures. Using these inputs, we design four interaction modes for different tasks, including Spherical Mode and Cartesian Mode for long-range moving, Operation Mode for precise manipulation, as well as Locking Mode for temporary pauses, where the hand gestures are utilized for seamless mode switching. We evaluate our system on a vertically mounted valve-turning task in the real world, demonstrating how each mode contributes to effective aerial manipulation. This interaction framework bridges human dexterity with aerial robotics, paving the way for enhanced aerial teleoperation in unstructured environments. Jinjie Li, Kotaro Kaneko, Haokun Liu, Liming Shu, Moju Zhao |
IROS | 4 |
| 2025 | Enhancing Training Data Attribution with Representational OptimizationabstractTraining data attribution (TDA) methods aim to measure how training data impacts a model's predictions.
While gradient-based attribution methods, such as influence functions, offer theoretical grounding, their computational costs make them impractical for large-scale applications.
Representation-based approaches are far more scalable, but typically rely on heuristic embeddings that are not optimized for attribution, limiting their fidelity.
To address these challenges, we propose AirRep,
a scalable, representation-based approach that closes this gap by learning task-specific and model-aligned representations optimized explicitly for TDA.
AirRep introduces two key innovations: a trainable encoder tuned for attribution quality, and an attention-based pooling mechanism that enables accurate estimation of group-wise influence.
We train AirRep using a ranking objective over automatically constructed training subsets labeled by their empirical effect on target predictions.
Experiments on instruction-tuned LLMs
demonstrate that AirRep achieves performance on par with state-of-the-art gradient-based approaches while being nearly two orders of magnitude more efficient at inference time.
Further analysis highlights its robustness
and generalization across tasks and models.
Our code is available at https://github.com/sunnweiwei/AirRep. Weiwei Sun 0001, Haokun Liu, Nikhil Kandpal, Colin Raffel, Yiming Yang 0002 |
NeurIPS | 2 |
| 2024 | Learning to Route Among Specialized Experts for Zero-Shot GeneralizationabstractRecently, there has been a widespread proliferation of "expert" language models that are specialized to a specific task or domain through parameter-efficient fine-tuning. How can we recycle large collections of expert language models to improve zero-shot generalization to unseen tasks? In this work, we propose $\textbf{P}$ost-$\textbf{H}$oc $\textbf{A}$daptive $\textbf{T}$okenwise $\textbf{G}$ating $\textbf{O}$ver an $\textbf{O}$cean of $\textbf{S}$pecialized $\textbf{E}$xperts (**PHATGOOSE**), which learns to route among specialized modules that were produced through parameter-efficient fine-tuning. Unlike past methods that learn to route among specialized models, PHATGOOSE explores the possibility that zero-shot generalization will be improved if different experts can be adaptively chosen for each token and at each layer in the model. Crucially, our method is *post-hoc* - it does not require simultaneous access to the datasets used to create the specialized models and only requires a modest amount of additional compute after each expert model is trained. In experiments covering a range of specialized model collections and zero-shot generalization benchmarks, we find that PHATGOOSE outperforms past methods for post-hoc routing and, in some cases, outperforms explicit multitask training (which requires simultaneous data access). To better understand the routing strategy learned by PHATGOOSE, we perform qualitative experiments to validate that PHATGOOSE's performance stems from its ability to make adaptive per-token and per-module expert choices. Mohammed Muqeeth, Haokun Liu, Colin Raffel |
ICML | 2 |
| 2024 | A fast intra CU partition algorithm in Versatile Video Coding for 360-degree video
Tiansong Li, Haokun Liu, Shaoguo Cui, Li Yu 0003, Kejun Wu, Hongkui Wang |
J. Vis. Commun. Image Represent. | 3 |
| 2024 | Fast CU partition algorithm based on swin-transformer for depth intra coding in 3D-HEVC
Shucen Liu, Shaoguo Cui, Tiansong Li, Haokun Liu, Qingsong Yang |
Multim. Tools Appl. | 4 |
| 2023 | Git-Theta: A Git Extension for Collaborative Development of Machine Learning ModelsabstractCurrently, most machine learning models are trained by centralized teams and are rarely updated. In contrast, open-source software development involves the iterative development of a shared artifact through distributed collaboration using a version control system. In the interest of enabling collaborative and continual improvement of machine learning models (Raffel, 2023), we introduce Git-Theta, a version control system for machine learning models. Git-Theta is an extension to Git, the most widely used version control software, that allows fine-grained tracking of changes to model parameters alongside code and other artifacts. Unlike existing version control systems that treat a model checkpoint as a blob of data, Git-Theta leverages the structure of checkpoints to support communication-efficient updates, automatic model merges, and meaningful reporting about the difference between two versions of a model. In addition, Git-Theta includes a plug-in system that enables users to easily add support for new functionality. In this paper, we introduce Git-Theta’s design and features and include an example use-case of Git-Theta where a pre-trained model is continually adapted and modified. We publicly release Git-Theta in hopes of kickstarting a new era of collaborative model development. https://github.com/r-three/git-theta/ Nikhil Kandpal, Brian Lester, Mohammed Muqeeth, Anisha Mascarenhas, Monty Evans, Vishal Baskaran, Tenghao Huang, Haokun Liu, Colin Raffel |
ICML | 8 |
| 2022 | Few-Shot Parameter-Efficient Fine-Tuning is Better and Cheaper than In-Context LearningabstractFew-shot in-context learning (ICL) enables pre-trained language models to perform a previously-unseen task without any gradient-based training by feeding a small number of training examples as part of the input. ICL incurs substantial computational, memory, and storage costs because it involves processing all of the training examples every time a prediction is made. Parameter-efficient fine-tuning (PEFT) (e.g. adapter modules, prompt tuning, sparse update methods, etc.) offers an alternative paradigm where a small set of parameters are trained to enable a model to perform the new task. In this paper, we rigorously compare few-shot ICL and PEFT and demonstrate that the latter offers better accuracy as well as dramatically lower computational costs. Along the way, we introduce a new PEFT method called (IA)^3 that scales activations by learned vectors, attaining stronger performance while only introducing a relatively tiny amount of new parameters. We also propose a simple recipe based on the T0 model called T-Few that can be applied to new tasks without task-specific tuning or modifications. We validate the effectiveness of T-Few on completely unseen tasks by applying it to the RAFT benchmark, attaining super-human performance for the first time and outperforming the state-of-the-art by 6% absolute. All of the code used in our experiments will be publicly available. Haokun Liu, Derek Tam, Mohammed Muqeeth, Jay Mohta, Tenghao Huang, Mohit Bansal, Colin Raffel |
NeurIPS | 1 |
| 2021 | Comparing Test Sets with Item Response TheoryabstractClara Vania, Phu Mon Htut, William Huang, Dhara Mungra, Richard Yuanzhe Pang, Jason Phang, Haokun Liu, Kyunghyun Cho, Samuel R. Bowman. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Clara Vania, Phu Mon Htut, William Huang, Dhara A. Mungra, Richard Yuanzhe Pang, Jason Phang, Haokun Liu, Kyunghyun Cho, Samuel R. Bowman |
ACL/IJCNLP (1) | 7 |
| 2020 | Intermediate-Task Transfer Learning with Pretrained Language Models: When and Why Does It Work?abstractYada Pruksachatkun, Jason Phang, Haokun Liu, Phu Mon Htut, Xiaoyi Zhang, Richard Yuanzhe Pang, Clara Vania, Katharina Kann, Samuel R. Bowman. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. Yada Pruksachatkun, Jason Phang, Haokun Liu, Phu Mon Htut, Richard Yuanzhe Pang, Clara Vania, Katharina Kann, Samuel R. Bowman |
ACL | 3 |
| 2020 | Precise Task Formalization Matters in Winograd Schema EvaluationsabstractPerformance on the Winograd Schema Challenge (WSC), a respected English commonsense reasoning benchmark, recently rocketed from chance accuracy to 89% on the Super-GLUE leaderboard, with relatively little corroborating evidence of a correspondingly large improvement in reasoning ability.We hypothesize that much of this improvement comes from recent changes in task formalizationthe combination of input specification, loss function, and reuse of pretrained parametersby users of the dataset, rather than improvements in the pretrained model's reasoning ability.We perform an ablation on two Winograd Schema datasets that interpolates between the formalizations used before and after this surge, and find (i) framing the task as multiple choice improves performance by 2-6 points and (ii) several additional techniques, including the reuse of a pretrained language modeling head, can mitigate the model's extreme sensitivity to hyperparameters.We urge future benchmark creators to impose additional structure to minimize the impact of formalization decisions on reported results. Haokun Liu, William Huang, Dhara A. Mungra, Samuel R. Bowman |
EMNLP (1) | 1 |
| 2020 | Learning Which Features Matter: RoBERTa Acquires a Preference for Linguistic Generalizations (Eventually)abstractOne reason pretraining on self-supervised linguistic tasks is effective is that it teaches models features that are helpful for language understanding.However, we want pretrained models to learn not only to represent linguistic features, but also to use those features preferentially during fine-turning.With this goal in mind, we introduce a new English-language diagnostic set called MSGS (the Mixed Signals Generalization Set), which consists of 20 ambiguous binary classification tasks that we use to test whether a pretrained model prefers linguistic or surface generalizations during finetuning.We pretrain RoBERTa models from scratch on quantities of data ranging from 1M to 1B words and compare their performance on MSGS to the publicly available RoBERTa BASE .We find that models can learn to represent linguistic features with little pretraining data, but require far more data to learn to prefer linguistic generalizations over surface ones.Eventually, with about 30B words of pretraining data, RoBERTa BASE does demonstrate a linguistic bias with some regularity.We conclude that while self-supervised pretraining is an effective way to learn helpful inductive biases, there is likely room to improve the rate at which models learn which features matter. Feature type Feature description Positive example Negative example SurfaceAbsolute position Is the first token of S "the"?The cat chased a mouse.A cat chased a mouse.Length Is S longer than n (e.g., 3) words?The cat chased a mouse.The cat meowed.Lexical content Does S contain "the"?That cat chased the mouse.That cat chased a mouse.Relative position Does "the" precede "a"?The cat chased a mouse.A cat chased the mouse.Orthography Does S appear in title case?The Cat Chased a Mouse.The cat chased a mouse. LinguisticMorphology Does S have an irregular past verb?The cats slept.The cats meow.Syn.category Does S have an adjective?Lincoln was tall.Lincoln was president.Syn.construction Is S the control construction?Sue is eager to sleep.Sue is likely to sleep.Syn.position Is the main verb in "ing" form?Cats who eat mice are purring.Cats who are eating mice purr. Alex Warstadt, Yian Zhang, Xiaocheng Li, Haokun Liu, Samuel R. Bowman |
EMNLP (1) | 4 |
| 2020 | BLiMP: The Benchmark of Linguistic Minimal Pairs for EnglishabstractWe introduce The Benchmark of Linguistic Minimal Pairs (BLiMP),1 a challenge set for evaluating the linguistic knowledge of language models (LMs) on major grammatical phenomena in English. BLiMP consists of 67 individual datasets, each containing 1,000 minimal pairs—that is, pairs of minimally different sentences that contrast in grammatical acceptability and isolate specific phenomenon in syntax, morphology, or semantics. We generate the data according to linguist-crafted grammar templates, and human aggregate agreement with the labels is 96.4%. We evaluate n-gram, LSTM, and Transformer (GPT-2 and Transformer-XL) LMs by observing whether they assign a higher probability to the acceptable sentence in each minimal pair. We find that state-of-the-art models identify morphological contrasts related to agreement reliably, but they struggle with some subtle semantic and syntactic phenomena, such as negative polarity items and extraction islands. Alex Warstadt, Alicia Parrish, Haokun Liu, Anhad Mohananey, Wei Peng 0013, Sheng-Fu Wang, Samuel R. Bowman |
Trans. Assoc. Comput. Linguistics | 3 |
| 2020 | Erratum: "BLiMP: The Benchmark of Linguistic Minimal Pairs for English"abstractWe correct wrongly reported results on BLiMP. Alex Warstadt, Alicia Parrish, Haokun Liu, Anhad Mohananey, Wei Peng 0013, Sheng-Fu Wang, Samuel R. Bowman |
Trans. Assoc. Comput. Linguistics | 3 |
| 2019 | Investigating BERT's Knowledge of Language: Five Analysis Methods with NPIsabstractAlex Warstadt, Yu Cao, Ioana Grosu, Wei Peng, Hagen Blix, Yining Nie, Anna Alsop, Shikha Bordia, Haokun Liu, Alicia Parrish, Sheng-Fu Wang, Jason Phang, Anhad Mohananey, Phu Mon Htut, Paloma Jeretic, Samuel R. Bowman. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Alex Warstadt, Ioana Grosu, Wei Peng 0013, Hagen Blix, Yining Nie, Anna Alsop, Shikha Bordia, Haokun Liu, Alicia Parrish, Sheng-Fu Wang, Jason Phang, Anhad Mohananey, Phu Mon Htut, Paloma Jeretic, Samuel R. Bowman |
EMNLP/IJCNLP (1) | 9 |
| 2018 | MEMD: A Diversity-Promoting Learning Framework for Short-Text ConversationabstractNeural encoder-decoder models have been widely applied to conversational response generation, which is a research hot spot in recent years. However, conventional neural encoder-decoder models tend to generate commonplace responses like “I don’t know” regardless of what the input is. In this paper, we analyze this problem from a new perspective: latent vectors. Based on it, we propose an easy-to-extend learning framework named MEMD (Multi-Encoder to Multi-Decoder), in which an auxiliary encoder and an auxiliary decoder are introduced to provide necessary training guidance without resorting to extra data or complicating network’s inner structure. Experimental results demonstrate that our method effectively improve the quality of generated responses according to automatic metrics and human evaluations, yielding more diverse and smooth replies. Meng Zou, Xihan Li 0001, Haokun Liu, Zhi-Hong Deng 0001 |
COLING | 3 |