Hao Lang

dblp:71/6934 · DBLP profile ↗
← Back
11ranked-venue papers
8as first author
8since 2021 · last 2026
0000-0002-6725-5898ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 6 first-author · 8 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Language models and text generation · 40% Question answering and dialogue systems · 24% Reinforcement learning · 17%

Topics — the 10 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
alignment
1.922026
Selective Weak-to-Strong Generalization · AAAI 2026
Debate Helps Weak-to-Strong Generalization · AAAI 2025
Natural language and speech › Language models and text generation › alignment › scalable oversight
weak-to-strong generalization
1.922026
Selective Weak-to-Strong Generalization · AAAI 2026
Debate Helps Weak-to-Strong Generalization · AAAI 2025
Machine learning › Reinforcement learning › function approximation › representation learning for reinforcement learning
latent action space
1.012026
Controlling Multimodal Conversational Agents with Coverage-Enhanced Latent Actions · ACL (1) 2026
Natural language and speech › Question answering and dialogue systems
multimodal dialogue system
1.012026
Controlling Multimodal Conversational Agents with Coverage-Enhanced Latent Actions · ACL (1) 2026
Machine learning › Reinforcement learning › reinforcement learning for NLP
reinforcement fine-tuning
1.012026
Controlling Multimodal Conversational Agents with Coverage-Enhanced Latent Actions · ACL (1) 2026
Natural language and speech › Language models and text generation › alignment
scalable oversight
0.912025
Debate Helps Weak-to-Strong Generalization · AAAI 2025
Natural language and speech › Question answering and dialogue systems
intent detection
0.612022
Estimating Soft Labels for Out-of-Domain Intent Detection · EMNLP 2022
Natural language and speech › Question answering and dialogue systems › intent detection
out-of-scope intent detection
0.612022
Estimating Soft Labels for Out-of-Domain Intent Detection · EMNLP 2022
Machine learning › Graph learning › graph signal processing
graph smoothing
0.312026
Selective Weak-to-Strong Generalization · AAAI 2026
Machine learning › Graph learning › limited supervision › multi-view semi-supervised learning
co-training
0.212022
Estimating Soft Labels for Out-of-Domain Intent Detection · EMNLP 2022

Methods — techniques the papers use, named apart from their topics

learning from observation · 1.0graph smoothing · 1.0cycle-consistency loss · 1.0cross-modal projector · 1.0binary classifier · 1.0fine-tuning · 0.9ensemble · 0.9retrieve-then-rerank · 0.7knowledge distillation · 0.7in-context learning · 0.7
YearPublicationVenuePosition
2026 Selective Weak-to-Strong Generalization
abstract
Future superhuman models will surpass the ability of humans and humans will only be able to \textit{weakly} supervise superhuman models. To alleviate the issue of lacking high-quality data for model alignment, some works on weak-to-strong generalization (W2SG) finetune a strong pretrained model with a weak supervisor so that it can generalize beyond weak supervision. However, the invariable use of weak supervision in existing methods exposes issues in robustness, with a proportion of weak labels proving harmful to models. In this paper, we propose a selective W2SG framework to avoid using weak supervision when unnecessary. We train a binary classifier P(IK) to identify questions that a strong model can answer and use its self-generated labels for alignment. We further refine weak labels with a graph smoothing method. Extensive experiments on three benchmarks show that our method consistently outperforms competitive baselines. Further analyses show that P(IK) can generalize across tasks and difficulties, which indicates selective W2SG can help superalignment.
Hao Lang, Fei Huang 0002, Yongbin Li 0001
AAAI1
2026 Controlling Multimodal Conversational Agents with Coverage-Enhanced Latent Actions
abstract
Vision-language models are increasingly employed as multimodal conversational agents (MCAs) for diverse conversational tasks.Recently, reinforcement learning (RL) has been widely explored for adapting MCAs to various human-AI interaction scenarios.Despite showing great enhancement in generalization performance, fine-tuning MCAs via RL still faces challenges in handling the extremely large text token space.To address this, we learn a compact latent action space for RL fine-tuning instead.Specifically, we adopt the learning from observation mechanism to construct the codebook for the latent action space, where future observations are leveraged to estimate current latent actions that could further be used to reconstruct future observations.However, the scarcity of paired image-text data hinders learning a codebook with sufficient coverage.Thus, we leverage both paired image-text data and text-only data to construct the latent action space, using a cross-modal projector for transforming text embeddings into image-text embeddings.We initialize the cross-modal projector on paired image-text data, and further train it on massive text-only data with a novel cycle consistency loss to enhance its robustness.We show that our latent action based method outperforms competitive baselines on two conversation tasks across various RL algorithms.
Yongqi Li 0002, Hao Lang, Tieyun Qian, Yongbin Li 0001
ACL (1)2
2025 Debate Helps Weak-to-Strong Generalization
abstract
Common methods for aligning already-capable models with desired behavior rely on the ability of humans to provide supervision. However, future superhuman models will surpass the capability of humans. Therefore, humans will only be able to weakly supervise superhuman models. This expected deficiency of human evaluation would weaken the safety of future AI systems. Scalable oversight and weak-to-strong generalization are two complementary approaches to tackle this issue. In this paper, we attempt to combine the strengths of these two approaches to further improve alignment. Specifically, we investigate ways of improving human supervision with a strong pretrained model and then supervise the strong model with enhanced weak human supervision. To make iterative empirical progress, we consider an analogy: can we use a strong model to improve weak model supervision and then use it to supervise the strong model? We empirically test it by finetuning a small weak model on ground truth labels with the additional help from a large strong model, and then finetuning the strong model on labels generated by the weak model. We find that debate can assist a weak model in extracting trustworthy information from an untrustworthy strong model, which provides leverage as context on samples when training a weak model. We also show that an ensemble of weak models helps exploit long arguments generated by strong model debaters and obtain a more robust supervision estimate. Extensive experiments on the OpenAI weak-to-strong NLP benchmarks show that the combination approach leads to better alignment, which indicates that debate has the potential to help weak-to-strong generalization.
Hao Lang, Fei Huang 0002, Yongbin Li 0001
AAAI1
2024 Scaling Data Diversity for Fine-Tuning Language Models in Human Alignment
abstract
Alignment with human preference prevents large language models (LLMs) from generating misleading or toxic content while requiring high-cost human feedback. Assuming resources of human annotation are limited, there are two different ways of allocating considered: more diverse PROMPTS or more diverse RESPONSES to be labeled. Nonetheless, a straightforward comparison between their impact is absent. In this work, we first control the diversity of both sides according to the number of samples for fine-tuning, which can directly reflect their influence. We find that instead of numerous prompts, more responses but fewer prompts better trigger LLMs for human alignment. Additionally, the concept of diversity for prompts can be more complex than responses that are typically quantified by single digits. Consequently, a new formulation of prompt diversity is proposed, further implying a linear correlation with the final performance of LLMs after fine-tuning. We also leverage it on data augmentation and conduct experiments to show its effect on different algorithms.
Feifan Song 0001, Bowen Yu 0002, Hao Lang, Haiyang Yu 0003, Fei Huang 0002, Houfeng Wang, Yongbin Li 0001
LREC/COLING3
2024 Out-of-Domain Intent Detection Considering Multi-Turn Dialogue Contexts
abstract
Out-of-Domain (OOD) intent detection is vital for practical dialogue systems, and it usually requires considering multi-turn dialogue contexts. However, most previous OOD intent detection approaches are limited to single dialogue turns. In this paper, we introduce a context-aware OOD intent detection (Caro) framework to model multi-turn contexts in OOD intent detection tasks. Specifically, we follow the information bottleneck principle to extract robust representations from multi-turn dialogue contexts. Two different views are constructed for each input sample and the superfluous information not related to intent detection is removed using a multi-view information bottleneck loss. Moreover, we also explore utilizing unlabeled data in Caro. A two-stage training process is introduced to mine OOD samples from these unlabeled data, and these OOD samples are used to train the resulting model with a bootstrapping approach. Comprehensive experiments demonstrate that Caro establishes state-of-the-art performances on multi-turn OOD detection tasks by improving the F1-OOD score of over 29% compared to the previous best method.
Hao Lang, Yinhe Zheng, Binyuan Hui, Fei Huang 0002, Yongbin Li 0001
LREC/COLING1
2024 Fine-Tuning Language Models with Reward Learning on Policy
abstract
Hao Lang, Fei Huang, Yongbin Li. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Hao Lang, Fei Huang 0002, Yongbin Li 0001
NAACL-HLT1
2023 Long-Tailed Question Answering in an Open World
abstract
Real-world data often have an open long-tailed distribution, and building a unified QA model supporting various tasks is vital for practical QA applications.However, it is non-trivial to extend previous QA approaches since they either require access to seen tasks of adequate samples or do not explicitly model samples from unseen tasks.In this paper, we define Open Long-Tailed QA (OLTQA) as learning from long-tailed distributed data and optimizing performance over seen and unseen QA tasks.We propose an OLTQA model that encourages knowledge sharing between head, tail and unseen tasks, and explicitly mines knowledge from a large pre-trained language model (LM).Specifically, we organize our model through a pool of fine-grained components and dynamically combine these components for an input to facilitate knowledge sharing.A retrieve-then-rerank frame is further introduced to select in-context examples, which guild the LM to generate text that express knowledge for QA tasks.Moreover, a twostage training approach is introduced to pretrain the framework by knowledge distillation (KD) from the LM and then jointly train the frame and a QA model through an adaptive mutual KD method.On a large-scale OLTQA dataset we curate from 43 existing QA datasets, our model consistently outperforms the stateof-the-art.We release the code and data
Hao Lang, Yinhe Zheng, Fei Huang 0002, Yongbin Li 0001
ACL (1)2
2022 Estimating Soft Labels for Out-of-Domain Intent Detection
abstract
Out-of-Domain (OOD) intent detection is important for practical dialog systems.To alleviate the issue of lacking OOD training samples, some works propose synthesizing pseudo OOD samples and directly assigning one-hot OOD labels to these pseudo samples.However, these one-hot labels introduce noises to the training process because some "hard" pseudo OOD samples may coincide with In-Domain (IND) intents.In this paper, we propose an adaptive soft pseudo labeling (ASoul) method that can estimate soft labels for pseudo OOD samples when training OOD detectors.Semantic connections between pseudo OOD samples and IND intents are captured using an embedding graph.A co-training framework is further introduced to produce resulting soft labels following the smoothness assumption, i.e., close samples are likely to have similar labels.Extensive experiments on three benchmark datasets show that ASoul consistently improves the OOD detection performance and outperforms various competitive baselines.
Hao Lang, Yinhe Zheng, Jian Sun 0021, Fei Huang 0002, Luo Si, Yongbin Li 0001
EMNLP1
2010 Improved latent concept expansion using hierarchical markov random fields
abstract
Most existing query expansion approaches for ad-hoc retrieval adopt overly simplistic textual representations that treat documents as bags of words and ignore inherent document structure. These simple representations often lead to incorrect independence assumptions in the proposed approaches and result in limited retrieval effectiveness. In this paper, we propose a novel query expansion technique that models the various types of dependencies that exist between original query terms and expansion terms within a robust, unified framework. The proposed model is called Hierarchical Markov random fields (HMRFs), based on Latent Concept Expansion (LCE). By exploiting implicit (or explicit) hierarchical structure within documents, HMRFs can incorporate hierarchical interactions which are important for modeling term dependencies in an efficient manner. Our rigorous experimental evaluation carried out using several TREC data sets shows that our proposed query expansion technique consistently and significantly outperforms the current state-of-the-art query expansion approaches, including relevance-based language models and LCE.
Hao Lang, Donald Metzler, Bin Wang 0004, Jintao Li 0001
CIKM1
2008 An Evaluation and Analysis of Incorporating Term Dependency for Ad-Hoc Retrieval
Hao Lang, Bin Wang 0004, Gareth J. F. Jones, Jintao Li 0001
ECIR1
2008 Query Performance Prediction for Information Retrieval Based on Covering Topic Score
Hao Lang, Bin Wang 0004, Gareth J. F. Jones, Jintao Li 0001
J. Comput. Sci. Technol.1