Haokun Liu

dblp:169/0460 · DBLP profile ↗
← Back
17ranked-venue papers
3as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
10 papers
Language models and text generation · 33% Efficient and distributed learning · 16% Trustworthy machine learning · 14%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Performance modeling and evaluation · 100%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational science and engineering · 100%
Software engineering, system software, and programming languages
1 paper
Requirements engineering and software design · 100%

Topics — the 21 heaviest of 25, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning
1.322024
Learning to Route Among Specialized Experts for Zero-Shot Generalization · ICML 2024
Few-Shot Parameter-Efficient Fine-Tuning is Better and Cheaper than In-Context Learning · NeurIPS 2022
Natural language and speech › Language models and text generation
pre-trained language model
0.932020
Learning Which Features Matter: RoBERTa Acquires a Preference for Linguistic Generalizations (Eventually) · EMNLP (1) 2020
Investigating BERT's Knowledge of Language: Five Analysis Methods with NPIs · EMNLP/IJCNLP (1) 2019
Intermediate-Task Transfer Learning with Pretrained Language Models: When and Why Does It Work? · ACL 2020
Computer vision › 3D vision › geometric estimation › geometric model fitting
hypothesis generation
0.912025
Literature Meets Data: A Synergistic Approach to Hypothesis Generation · ACL (1) 2025
Machine learning › Trustworthy machine learning
interpretability
0.912025
Enhancing Training Data Attribution with Representational Optimization · NeurIPS 2025
Machine learning › Trustworthy machine learning › interpretability
training data attribution
0.912025
Enhancing Training Data Attribution with Representational Optimization · NeurIPS 2025
Computational science and engineering › AI for science
AI for scientific discovery
0.912025
Literature Meets Data: A Synergistic Approach to Hypothesis Generation · ACL (1) 2025
Natural language and speech › Language models and text generation
linguistic generalization
0.822020
Learning Which Features Matter: RoBERTa Acquires a Preference for Linguistic Generalizations (Eventually) · EMNLP (1) 2020
Investigating BERT's Knowledge of Language: Five Analysis Methods with NPIs · EMNLP/IJCNLP (1) 2019
Machine learning › Deep learning architectures and training › mixture of experts
expert routing
0.812024
Learning to Route Among Specialized Experts for Zero-Shot Generalization · ICML 2024
Machine learning › Transfer learning and domain adaptation
zero-shot transfer
0.812024
Learning to Route Among Specialized Experts for Zero-Shot Generalization · ICML 2024
Requirements engineering and software design › model-driven engineering
collaborative modeling
0.712023
Git-Theta: A Git Extension for Collaborative Development of Machine Learning Models · ICML 2023
Natural language and speech › Language models and text generation
in-context learning
0.612022
Few-Shot Parameter-Efficient Fine-Tuning is Better and Cheaper than In-Context Learning · NeurIPS 2022
Natural language and speech › Language models and text generation › large language model evaluation
NLP evaluation
0.512021
Comparing Test Sets with Item Response Theory · ACL/IJCNLP (1) 2021
Performance modeling and evaluation
benchmarking
0.512021
Comparing Test Sets with Item Response Theory · ACL/IJCNLP (1) 2021
Performance modeling and evaluation › statistical analysis
item response theory
0.512021
Comparing Test Sets with Item Response Theory · ACL/IJCNLP (1) 2021
Knowledge, reasoning and agents › Knowledge representation and reasoning
commonsense reasoning
0.412020
Precise Task Formalization Matters in Winograd Schema Evaluations · EMNLP (1) 2020
Machine learning › Learning theory
inductive bias
0.412020
Learning Which Features Matter: RoBERTa Acquires a Preference for Linguistic Generalizations (Eventually) · EMNLP (1) 2020
Natural language and speech › Language models and text generation › pre-trained language model
RoBERTa
0.412020
Learning Which Features Matter: RoBERTa Acquires a Preference for Linguistic Generalizations (Eventually) · EMNLP (1) 2020
Knowledge, reasoning and agents › Knowledge representation and reasoning › commonsense reasoning
winograd schema challenge
0.412020
Precise Task Formalization Matters in Winograd Schema Evaluations · EMNLP (1) 2020
Natural language and speech › Language models and text generation › pre-trained language model › knowledge probing
linguistic knowledge probing
0.412019
Investigating BERT's Knowledge of Language: Five Analysis Methods with NPIs · EMNLP/IJCNLP (1) 2019
Natural language and speech › Information extraction and text analysis › text classification
deception detection
0.312025
Literature Meets Data: A Synergistic Approach to Hypothesis Generation · ACL (1) 2025
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
self-supervised representation learning
0.112020
Learning Which Features Matter: RoBERTa Acquires a Preference for Linguistic Generalizations (Eventually) · EMNLP (1) 2020

Methods — techniques the papers use, named apart from their topics

literature-based retrieval · 1.7data-driven generation · 1.7version control · 1.3model merging · 1.3ranking objective · 0.9influence functions · 0.9fine-tuning · 0.9attention-based pooling · 0.9tokenwise gating · 0.8parameter-efficient fine-tuning · 0.8item response theory · 0.5
YearPublicationVenuePosition
2026 Hierarchical Trajectory Planning of Floating-Base Multi-Link Robot for Maneuvering in Confined Environments
abstract
Floating-base multi-link robots can change their shape during flight, making them well-suited for applications in confined environments such as autonomous inspection and search and rescue. However, trajectory planning for such systems remains an open challenge because the problem lies in a high-dimensional, constraint-rich space where collision avoidance must be addressed together with kinematic limits and dynamic feasibility. This work introduces a hierarchical trajectory planning framework that integrates global guidance with configuration-aware local optimization. First, we exploit the dual nature of these robots—the root link as a rigid body for guidance and the articulated joints for flexibility—to generate global anchor states that decompose the planning problem into tractable segments. Second, we design a local trajectory planner that optimizes each segment in parallel with differentiable objectives and constraints, systematically enforcing kinematic feasibility and maintaining dynamic feasibility by avoiding control singularities. Third, we implement a complete system that directly processes point-cloud data, eliminating the need for handcrafted obstacle models. Extensive simulations and real-world experiments confirm that this framework enables an articulated aerial robot to exploit its morphology for maneuvering that rigid robots cannot achieve. To the best of our knowledge, this is the first planning framework for floating-base multi-link robots that has been demonstrated on a real robot to generate continuous, collision-free, and dynamically feasible trajectories directly from raw point-cloud inputs, without relying on handcrafted obstacle models.
Jinjie Li, Haokun Liu, Zicheng Luo, Kotaro Kaneko, Moju Zhao
IEEE Trans Autom. Sci. Eng.3
2025 Literature Meets Data: A Synergistic Approach to Hypothesis Generation
abstract
AI holds promise for transforming scientific processes, including hypothesis generation.Prior work on hypothesis generation can be broadly categorized into theory-driven and datadriven approaches.While both have proven effective in generating novel and plausible hypotheses, it remains an open question whether they can complement each other.To address this, we develop the first method that combines literature-based insights with data to perform LLM-powered hypothesis generation.We apply our method on five different datasets and demonstrate that integrating literature and data outperforms other baselines (8.97% over fewshot, 15.75% over literature-based alone, and 3.37% over data-driven alone).Additionally, we conduct the first human evaluation to assess the utility of LLM-generated hypotheses in assisting human decision-making on two challenging tasks: deception detection and AI generated content detection.Our results show that human accuracy improves significantly by 7.44% and 14.19% on these tasks, respectively.These findings suggest that integrating literature-based and data-driven approaches provides a comprehensive and nuanced framework for hypothesis generation and could open new avenues for scientific inquiry.
Haokun Liu, Yangqiaoyu Zhou, Chenfei Yuan, Chenhao Tan
ACL (1)1
2025 Six-DoF Hand-Based Teleoperation for Omnidirectional Aerial Robots
abstract
Omnidirectional aerial robots offer full 6-DoF independent control over position and orientation, making them popular for aerial manipulation. Although advancements in robotic autonomy, human operation remains essential in complex aerial environments. Existing teleoperation approaches for multirotors fail to fully leverage the additional DoFs provided by omnidirectional rotation. Additionally, the dexterity of human fingers should be exploited for more engaged interaction. In this work, we propose an aerial teleoperation system that brings the rotational flexibility of human hands into the unbounded aerial workspace. Our system includes two motion-tracking marker sets—one on the shoulder and one on the hand—along with a data glove to capture hand gestures. Using these inputs, we design four interaction modes for different tasks, including Spherical Mode and Cartesian Mode for long-range moving, Operation Mode for precise manipulation, as well as Locking Mode for temporary pauses, where the hand gestures are utilized for seamless mode switching. We evaluate our system on a vertically mounted valve-turning task in the real world, demonstrating how each mode contributes to effective aerial manipulation. This interaction framework bridges human dexterity with aerial robotics, paving the way for enhanced aerial teleoperation in unstructured environments.
Jinjie Li, Kotaro Kaneko, Haokun Liu, Liming Shu, Moju Zhao
IROS4
2025 Enhancing Training Data Attribution with Representational Optimization
abstract
Training data attribution (TDA) methods aim to measure how training data impacts a model's predictions. While gradient-based attribution methods, such as influence functions, offer theoretical grounding, their computational costs make them impractical for large-scale applications. Representation-based approaches are far more scalable, but typically rely on heuristic embeddings that are not optimized for attribution, limiting their fidelity. To address these challenges, we propose AirRep, a scalable, representation-based approach that closes this gap by learning task-specific and model-aligned representations optimized explicitly for TDA. AirRep introduces two key innovations: a trainable encoder tuned for attribution quality, and an attention-based pooling mechanism that enables accurate estimation of group-wise influence. We train AirRep using a ranking objective over automatically constructed training subsets labeled by their empirical effect on target predictions. Experiments on instruction-tuned LLMs demonstrate that AirRep achieves performance on par with state-of-the-art gradient-based approaches while being nearly two orders of magnitude more efficient at inference time. Further analysis highlights its robustness and generalization across tasks and models. Our code is available at https://github.com/sunnweiwei/AirRep.
Weiwei Sun 0001, Haokun Liu, Nikhil Kandpal, Colin Raffel, Yiming Yang 0002
NeurIPS2
2024 Learning to Route Among Specialized Experts for Zero-Shot Generalization
abstract
Recently, there has been a widespread proliferation of "expert" language models that are specialized to a specific task or domain through parameter-efficient fine-tuning. How can we recycle large collections of expert language models to improve zero-shot generalization to unseen tasks? In this work, we propose $\textbf{P}$ost-$\textbf{H}$oc $\textbf{A}$daptive $\textbf{T}$okenwise $\textbf{G}$ating $\textbf{O}$ver an $\textbf{O}$cean of $\textbf{S}$pecialized $\textbf{E}$xperts (**PHATGOOSE**), which learns to route among specialized modules that were produced through parameter-efficient fine-tuning. Unlike past methods that learn to route among specialized models, PHATGOOSE explores the possibility that zero-shot generalization will be improved if different experts can be adaptively chosen for each token and at each layer in the model. Crucially, our method is *post-hoc* - it does not require simultaneous access to the datasets used to create the specialized models and only requires a modest amount of additional compute after each expert model is trained. In experiments covering a range of specialized model collections and zero-shot generalization benchmarks, we find that PHATGOOSE outperforms past methods for post-hoc routing and, in some cases, outperforms explicit multitask training (which requires simultaneous data access). To better understand the routing strategy learned by PHATGOOSE, we perform qualitative experiments to validate that PHATGOOSE's performance stems from its ability to make adaptive per-token and per-module expert choices.
Mohammed Muqeeth, Haokun Liu, Colin Raffel
ICML2
2024 A fast intra CU partition algorithm in Versatile Video Coding for 360-degree video
Tiansong Li, Haokun Liu, Shaoguo Cui, Li Yu 0003, Kejun Wu, Hongkui Wang
J. Vis. Commun. Image Represent.3
2024 Fast CU partition algorithm based on swin-transformer for depth intra coding in 3D-HEVC
Shucen Liu, Shaoguo Cui, Tiansong Li, Haokun Liu, Qingsong Yang
Multim. Tools Appl.4
2023 Git-Theta: A Git Extension for Collaborative Development of Machine Learning Models
abstract
Currently, most machine learning models are trained by centralized teams and are rarely updated. In contrast, open-source software development involves the iterative development of a shared artifact through distributed collaboration using a version control system. In the interest of enabling collaborative and continual improvement of machine learning models (Raffel, 2023), we introduce Git-Theta, a version control system for machine learning models. Git-Theta is an extension to Git, the most widely used version control software, that allows fine-grained tracking of changes to model parameters alongside code and other artifacts. Unlike existing version control systems that treat a model checkpoint as a blob of data, Git-Theta leverages the structure of checkpoints to support communication-efficient updates, automatic model merges, and meaningful reporting about the difference between two versions of a model. In addition, Git-Theta includes a plug-in system that enables users to easily add support for new functionality. In this paper, we introduce Git-Theta’s design and features and include an example use-case of Git-Theta where a pre-trained model is continually adapted and modified. We publicly release Git-Theta in hopes of kickstarting a new era of collaborative model development. https://github.com/r-three/git-theta/
Nikhil Kandpal, Brian Lester, Mohammed Muqeeth, Anisha Mascarenhas, Monty Evans, Vishal Baskaran, Tenghao Huang, Haokun Liu, Colin Raffel
ICML8
2022 Few-Shot Parameter-Efficient Fine-Tuning is Better and Cheaper than In-Context Learning
abstract
Few-shot in-context learning (ICL) enables pre-trained language models to perform a previously-unseen task without any gradient-based training by feeding a small number of training examples as part of the input. ICL incurs substantial computational, memory, and storage costs because it involves processing all of the training examples every time a prediction is made. Parameter-efficient fine-tuning (PEFT) (e.g. adapter modules, prompt tuning, sparse update methods, etc.) offers an alternative paradigm where a small set of parameters are trained to enable a model to perform the new task. In this paper, we rigorously compare few-shot ICL and PEFT and demonstrate that the latter offers better accuracy as well as dramatically lower computational costs. Along the way, we introduce a new PEFT method called (IA)^3 that scales activations by learned vectors, attaining stronger performance while only introducing a relatively tiny amount of new parameters. We also propose a simple recipe based on the T0 model called T-Few that can be applied to new tasks without task-specific tuning or modifications. We validate the effectiveness of T-Few on completely unseen tasks by applying it to the RAFT benchmark, attaining super-human performance for the first time and outperforming the state-of-the-art by 6% absolute. All of the code used in our experiments will be publicly available.
Haokun Liu, Derek Tam, Mohammed Muqeeth, Jay Mohta, Tenghao Huang, Mohit Bansal, Colin Raffel
NeurIPS1
2021 Comparing Test Sets with Item Response Theory
abstract
Clara Vania, Phu Mon Htut, William Huang, Dhara Mungra, Richard Yuanzhe Pang, Jason Phang, Haokun Liu, Kyunghyun Cho, Samuel R. Bowman. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Clara Vania, Phu Mon Htut, William Huang, Dhara A. Mungra, Richard Yuanzhe Pang, Jason Phang, Haokun Liu, Kyunghyun Cho, Samuel R. Bowman
ACL/IJCNLP (1)7
2020 Intermediate-Task Transfer Learning with Pretrained Language Models: When and Why Does It Work?
abstract
Yada Pruksachatkun, Jason Phang, Haokun Liu, Phu Mon Htut, Xiaoyi Zhang, Richard Yuanzhe Pang, Clara Vania, Katharina Kann, Samuel R. Bowman. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020.
Yada Pruksachatkun, Jason Phang, Haokun Liu, Phu Mon Htut, Richard Yuanzhe Pang, Clara Vania, Katharina Kann, Samuel R. Bowman
ACL3
2020 Precise Task Formalization Matters in Winograd Schema Evaluations
abstract
Performance on the Winograd Schema Challenge (WSC), a respected English commonsense reasoning benchmark, recently rocketed from chance accuracy to 89% on the Super-GLUE leaderboard, with relatively little corroborating evidence of a correspondingly large improvement in reasoning ability.We hypothesize that much of this improvement comes from recent changes in task formalizationthe combination of input specification, loss function, and reuse of pretrained parametersby users of the dataset, rather than improvements in the pretrained model's reasoning ability.We perform an ablation on two Winograd Schema datasets that interpolates between the formalizations used before and after this surge, and find (i) framing the task as multiple choice improves performance by 2-6 points and (ii) several additional techniques, including the reuse of a pretrained language modeling head, can mitigate the model's extreme sensitivity to hyperparameters.We urge future benchmark creators to impose additional structure to minimize the impact of formalization decisions on reported results.
Haokun Liu, William Huang, Dhara A. Mungra, Samuel R. Bowman
EMNLP (1)1
2020 Learning Which Features Matter: RoBERTa Acquires a Preference for Linguistic Generalizations (Eventually)
abstract
One reason pretraining on self-supervised linguistic tasks is effective is that it teaches models features that are helpful for language understanding.However, we want pretrained models to learn not only to represent linguistic features, but also to use those features preferentially during fine-turning.With this goal in mind, we introduce a new English-language diagnostic set called MSGS (the Mixed Signals Generalization Set), which consists of 20 ambiguous binary classification tasks that we use to test whether a pretrained model prefers linguistic or surface generalizations during finetuning.We pretrain RoBERTa models from scratch on quantities of data ranging from 1M to 1B words and compare their performance on MSGS to the publicly available RoBERTa BASE .We find that models can learn to represent linguistic features with little pretraining data, but require far more data to learn to prefer linguistic generalizations over surface ones.Eventually, with about 30B words of pretraining data, RoBERTa BASE does demonstrate a linguistic bias with some regularity.We conclude that while self-supervised pretraining is an effective way to learn helpful inductive biases, there is likely room to improve the rate at which models learn which features matter. Feature type Feature description Positive example Negative example SurfaceAbsolute position Is the first token of S "the"?The cat chased a mouse.A cat chased a mouse.Length Is S longer than n (e.g., 3) words?The cat chased a mouse.The cat meowed.Lexical content Does S contain "the"?That cat chased the mouse.That cat chased a mouse.Relative position Does "the" precede "a"?The cat chased a mouse.A cat chased the mouse.Orthography Does S appear in title case?The Cat Chased a Mouse.The cat chased a mouse. LinguisticMorphology Does S have an irregular past verb?The cats slept.The cats meow.Syn.category Does S have an adjective?Lincoln was tall.Lincoln was president.Syn.construction Is S the control construction?Sue is eager to sleep.Sue is likely to sleep.Syn.position Is the main verb in "ing" form?Cats who eat mice are purring.Cats who are eating mice purr.
Alex Warstadt, Yian Zhang, Xiaocheng Li, Haokun Liu, Samuel R. Bowman
EMNLP (1)4
2020 BLiMP: The Benchmark of Linguistic Minimal Pairs for English
abstract
We introduce The Benchmark of Linguistic Minimal Pairs (BLiMP),1 a challenge set for evaluating the linguistic knowledge of language models (LMs) on major grammatical phenomena in English. BLiMP consists of 67 individual datasets, each containing 1,000 minimal pairs—that is, pairs of minimally different sentences that contrast in grammatical acceptability and isolate specific phenomenon in syntax, morphology, or semantics. We generate the data according to linguist-crafted grammar templates, and human aggregate agreement with the labels is 96.4%. We evaluate n-gram, LSTM, and Transformer (GPT-2 and Transformer-XL) LMs by observing whether they assign a higher probability to the acceptable sentence in each minimal pair. We find that state-of-the-art models identify morphological contrasts related to agreement reliably, but they struggle with some subtle semantic and syntactic phenomena, such as negative polarity items and extraction islands.
Alex Warstadt, Alicia Parrish, Haokun Liu, Anhad Mohananey, Wei Peng 0013, Sheng-Fu Wang, Samuel R. Bowman
Trans. Assoc. Comput. Linguistics3
2020 Erratum: "BLiMP: The Benchmark of Linguistic Minimal Pairs for English"
abstract
We correct wrongly reported results on BLiMP.
Alex Warstadt, Alicia Parrish, Haokun Liu, Anhad Mohananey, Wei Peng 0013, Sheng-Fu Wang, Samuel R. Bowman
Trans. Assoc. Comput. Linguistics3
2019 Investigating BERT's Knowledge of Language: Five Analysis Methods with NPIs
abstract
Alex Warstadt, Yu Cao, Ioana Grosu, Wei Peng, Hagen Blix, Yining Nie, Anna Alsop, Shikha Bordia, Haokun Liu, Alicia Parrish, Sheng-Fu Wang, Jason Phang, Anhad Mohananey, Phu Mon Htut, Paloma Jeretic, Samuel R. Bowman. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Alex Warstadt, Ioana Grosu, Wei Peng 0013, Hagen Blix, Yining Nie, Anna Alsop, Shikha Bordia, Haokun Liu, Alicia Parrish, Sheng-Fu Wang, Jason Phang, Anhad Mohananey, Phu Mon Htut, Paloma Jeretic, Samuel R. Bowman
EMNLP/IJCNLP (1)9
2018 MEMD: A Diversity-Promoting Learning Framework for Short-Text Conversation
abstract
Neural encoder-decoder models have been widely applied to conversational response generation, which is a research hot spot in recent years. However, conventional neural encoder-decoder models tend to generate commonplace responses like “I don’t know” regardless of what the input is. In this paper, we analyze this problem from a new perspective: latent vectors. Based on it, we propose an easy-to-extend learning framework named MEMD (Multi-Encoder to Multi-Decoder), in which an auxiliary encoder and an auxiliary decoder are introduced to provide necessary training guidance without resorting to extra data or complicating network’s inner structure. Experimental results demonstrate that our method effectively improve the quality of generated responses according to automatic metrics and human evaluations, yielding more diverse and smooth replies.
Meng Zou, Xihan Li 0001, Haokun Liu, Zhi-Hong Deng 0001
COLING3