Haowei Lin

dblp:235/2798 · DBLP profile ↗
← Back
21ranked-venue papers
6as first author
19since 2021 · last 2026
0009-0006-9809-4835ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 5 first-author · 18 since 2021Systems, architecture and hardware · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Fine-grained alignment in medical pathology vision-language models via variational distillation
abstract
Pre-trained vision-language (V-L) models exhibit promising performance across various general-domain tasks. However, they fall short in medical pathology due to the critical need for fine-grained semantic alignment, which is essential for distinguishing subtle visual patterns across categories. This limitation is not merely due to domain gaps but stems from the inability to capture detailed, pathology-specific semantics. Previous efforts leveraging large language models (LLMs) or cross-modal training often introduce redundant or ambiguous cues, ultimately weakening generalization. To explicitly enhance fine-grained alignment, we propose a Variational Distillation framework tailored for Medical Pathology V-L models. This method introduces a dual-loop optimization mechanism that jointly distills and aligns semantic signals from both textual inputs and external LLM knowledge. Specifically, we use variational latent distributions to model semantic ambiguity and apply a KL-based loss to reduce differences between signals. This encourages the model to retain robust and generalizable features, enabling improved sensitivity to subtle semantic variations critical in pathology image understanding. During cross-modal alignment, the proposed method further amplifies modality-shared semantics while suppressing modality-specific noise and task-irrelevant factors, yielding more precise and pathology-aware image-text matching. Extensive experiments on five pathology benchmarks across three settings, including class generalization, few-shot learning, and cross-organ transfer, demonstrate that the proposed method consistently outperforms the existing approaches.
Runlin Huang, Haowei Lin, Weipeng Zhuo, Yiu-Ming Cheung, Hongmin Cai, Weifeng Su
Pattern Recognit.2
2025 Efficient and Asymptotically Unbiased Constrained Decoding for Large Language Models
abstract
In real-world applications of large language models, outputs are often required to be confined: selecting items from predefined product or document sets, generating phrases that comply with safety standards, or conforming to specialized formatting styles. To control the generation, constrained decoding has been widely adopted. However, existing prefix-tree-based constrained decoding is inefficient under GPU-based model inference paradigms, and it introduces unintended biases into the output distribution. This paper introduces Dynamic Importance Sampling for Constrained Decoding (DISC) with GPU-based Parallel Prefix-Verification (PPV), a novel algorithm that leverages dynamic importance sampling to achieve theoretically guaranteed asymptotic unbiasedness and overcomes the inefficiency of prefix-tree. Extensive experiments demonstrate the superiority of our method over existing methods in both efficiency and output quality. These results highlight the potential of our methods to improve constrained generation in applications where adherence to specific constraints is essential.
Haotian Ye, Himanshu Jain, Chong You, Ananda Theertha Suresh, Haowei Lin, James Zou 0001, Felix X. Yu
AISTATS5
2025 GROOT-2: Weakly Supervised Multimodal Instruction Following Agents
abstract
Developing agents that can follow multimodal instructions remains a fundamental challenge in robotics and AI. Although large-scale pre-training on unlabeled datasets has enabled agents to learn diverse behaviors, these agents often struggle with following instructions. While augmenting the dataset with instruction labels can mitigate this issue, acquiring such high-quality annotations at scale is impractical. To address this issue, we frame the problem as a semi-supervised learning task and introduce \agent, a multimodal instructable agent trained using a novel approach that combines weak supervision with latent variable models. Our method consists of two key components: constrained self-imitating, which utilizes large amounts of unlabeled demonstrations to enable the policy to learn diverse behaviors, and human intention alignment, which uses a smaller set of labeled demonstrations to ensure the latent space reflects human intentions. \agent’s effectiveness is validated across four diverse environments, ranging from video games to robotic manipulation, demonstrating its robust multimodal instruction-following capabilities.
Shaofei Cai, Bowei Zhang 0007, Haowei Lin, Xiaojian Ma 0001, Anji Liu, Yitao Liang
ICLR4
2025 TFG-Flow: Training-free Guidance in Multimodal Generative Flow
abstract
Given an unconditional generative model and a predictor for a target property (e.g., a classifier), the goal of training-free guidance is to generate samples with desirable target properties without additional training. As a highly efficient technique for steering generative models toward flexible outcomes, training-free guidance has gained increasing attention in diffusion models. However, existing methods only handle data in continuous spaces, while many scientific applications involve both continuous and discrete data (referred to as multimodality). Another emerging trend is the growing use of the simple and general flow matching framework in building generative foundation models, where guided generation remains under-explored. To address this, we introduce TFG-Flow, a novel training-free guidance method for multimodal generative flow. TFG-Flow addresses the curse-of-dimensionality while maintaining the property of unbiased sampling in guiding discrete variables. We validate TFG-Flow on four molecular design tasks and show that TFG-Flow has great potential in drug design by generating molecules with desired properties.
Haowei Lin, Shanda Li, Haotian Ye, Stefano Ermon, Yitao Liang, Jianzhu Ma
ICLR1
2025 Integrating Protein Dynamics into Structure-Based Drug Design via Full-Atom Stochastic Flows
abstract
The dynamic nature of proteins, influenced by ligand interactions, is essential for comprehending protein function and progressing drug discovery. Traditional structure-based drug design (SBDD) approaches typically target binding sites with rigid structures, limiting their practical application in drug development. While molecular dynamics simulation can theoretically capture all the biologically relevant conformations, the transition rate is dictated by the intrinsic energy barrier between them, making the sampling process computationally expensive. To overcome the aforementioned challenges, we propose to use generative modeling for SBDD considering conformational changes of protein pockets. We curate a dataset of apo and multiple holo states of protein-ligand complexes, simulated by molecular dynamics, and propose a full-atom flow model (and a stochastic version), named DynamicFlow, that learns to transform apo pockets and noisy ligands into holo pockets and corresponding 3D ligand molecules. Our method uncovers promising ligand molecules and corresponding holo conformations of pockets. Additionally, the resultant holo-like states provide superior inputs for traditional SBDD approaches, playing a significant role in practical drug discovery.
Xiangxin Zhou, Haowei Lin, Xinheng He, Jiaqi Guan, Yang Wang 0103, Qiang Liu 0006, Liang Wang 0001, Jianzhu Ma
ICLR3
2025 MCU: An Evaluation Framework for Open-Ended Game Agents
abstract
Developing AI agents capable of interacting with open-world environments to solve diverse tasks is a compelling challenge. However, evaluating such open-ended agents remains difficult, with current benchmarks facing scalability limitations. To address this, we introduce \textit{Minecraft Universe} (MCU), a comprehensive evaluation framework set within the open-world video game Minecraft. MCU incorporates three key components: (1) an expanding collection of 3,452 composable atomic tasks that encompasses 11 major categories and 41 subcategories of challenges; (2) a task composition mechanism capable of generating infinite diverse tasks with varying difficulty; and (3) a general evaluation framework that achieves 91.5\% alignment with human ratings for open-ended task assessment. Empirical results reveal that even state-of-the-art foundation agents struggle with the increasing diversity and complexity of tasks. These findings highlight the necessity of MCU as a robust benchmark to drive progress in AI agent development within open-ended environments. Our evaluation code and scripts are available at https://github.com/CraftJarvis/MCU.
Xinyue Zheng, Haowei Lin, Kaichen He, Qiang Fu 0016, Haobo Fu, Zilong Zheng, Yitao Liang
ICML2
2025 ADA: An Adaptive Augmentation Framework for Single-Source Domain Generalization in Medical Image Segmentation
Runlin Huang, Hongmin Cai, Weipeng Zhuo, Shangyan Cai, Haowei Lin, Wentao Fan 0001, Weifeng Su
MICCAI (10)5
2025 Neuro-dynamic programming-based event-triggered fault tolerant control for nonlinear systems with multiple faults
Haowei Lin, Weifeng Su, Runlin Huang, Bo Zhao 0015, Wentao Fan 0001
Neural Networks1
2025 JARVIS-1: Open-World Multi-Task Agents With Memory-Augmented Multimodal Language Models
abstract
Achieving human-like planning and control with multimodal observations in an open world is a key milestone for more functional generalist agents. Existing approaches can handle certain long-horizon tasks in an open world. However, they still struggle when the number of open-world tasks could potentially be infinite and lack the capability to progressively enhance task completion as game time progresses. We introduceJARVIS-1, an open-world agent that can perceive multimodal input (visual observations and human instructions), generate sophisticated plans, and perform embodied control, all within the popular yet challenging open-world Minecraft universe. Specifically, we developJARVIS-1 on top of pre-trained multimodal language models, which map visual observations and textual instructions to plans. The plans will be ultimately dispatched to the goal-conditioned controllers. We outfitJARVIS-1 with a multimodal memory, which facilitates planning using both pre-trained knowledge and its actual game survival experiences.JARVIS-1 is the existing most general agent in Minecraft, capable of completing over 200 different tasks using control and observation space similar to humans. These tasks range from short-horizon tasks, e.g., “chopping trees” to long-horizon ones, e.g., “obtaining a diamond pickaxe”.JARVIS-1 performs exceptionally well in short-horizon tasks, achieving nearly perfect performance. In the classic long-term task ofObtainDiamondPickaxe,JARVIS-1 surpasses the reliability of current state-of-the-art agents by 5 times and can successfully complete longer-horizon and more challenging tasks. Furthermore, we show thatJARVIS-1 is able toself-improvefollowing a life-long learning paradigm thanks to multimodal memory, sparking a more general intelligence and improved autonomy.
Shaofei Cai, Anji Liu, Yonggang Jin, Jinbing Hou, Bowei Zhang 0007, Haowei Lin, Zhaofeng He 0001, Zilong Zheng, Yaodong Yang 0001, Xiaojian Ma 0001, Yitao Liang
IEEE Trans. Pattern Anal. Mach. Intell.7
2024 Class Incremental Learning via Likelihood Ratio Based Task Prediction
abstract
Class incremental learning (CIL) is a challenging setting of continual learning, which learns a series of tasks sequentially. Each task consists of a set of unique classes. The key feature of CIL is that no task identifier (or task-id) is provided at test time. Predicting the task-id for each test sample is a challenging problem. An emerging theory-guided approach (called TIL+OOD) is to train a task-specific model for each task in a shared network for all tasks based on a task-incremental learning (TIL) method to deal with catastrophic forgetting. The model for each task is an out-of-distribution (OOD) detector rather than a conventional classifier. The OOD detector can perform both within-task (in-distribution (IND)) class prediction and OOD detection. The OOD detection capability is the key to task-id prediction during inference. However, this paper argues that using a traditional OOD detector for task-id prediction is sub-optimal because additional information (e.g., the replay data and the learned tasks) available in CIL can be exploited to design a better and principled method for task-id prediction. We call the new method TPL (Task-id Prediction based on Likelihood Ratio). TPL markedly outperforms strong CIL baselines and has negligible catastrophic forgetting. The code of TPL is publicly available at https://github.com/linhaowei1/TPL.
Haowei Lin, Yijia Shao, Weinan Qian, Ningxin Pan, Yiduo Guo, Bing Liu 0001
ICLR1
2024 Selecting Large Language Model to Fine-tune via Rectified Scaling Law
abstract
The ever-growing ecosystem of LLMs has posed a challenge in selecting the most appropriate pre-trained model to fine-tune amidst a sea of options. Given constrained resources, fine-tuning all models and making selections afterward is unrealistic. In this work, we formulate this resource-constrained selection task into predicting fine-tuning performance and illustrate its natural connection with Scaling Law. Unlike pre-training, we find that the fine-tuning scaling curve includes not just the well-known "power phase" but also the previously unobserved "pre-power phase". We also explain why existing Scaling Law fails to capture this phase transition phenomenon both theoretically and empirically. To address this, we introduce the concept of "pre-learned data size" into our Rectified Scaling Law, which overcomes theoretical limitations and fits experimental results much better. By leveraging our law, we propose a novel LLM selection algorithm that selects the near-optimal model with hundreds of times less resource consumption, while other methods may provide negatively correlated selection. The project page is available at rectified-scaling-law.github.io.
Haowei Lin, Baizhou Huang, Haotian Ye, Qinyu Chen, Sujian Li, Jianzhu Ma, Xiaojun Wan 0001, James Zou 0001, Yitao Liang
ICML1
2024 OmniJARVIS: Unified Vision-Language-Action Tokenization Enables Open-World Instruction Following Agents
abstract
This paper presents OmniJARVIS, a novel Vision-Language-Action (VLA) model for open-world instruction-following agents in Minecraft. Compared to prior works that either emit textual goals to separate controllers or produce the control command directly, OmniJARVIS seeks a different path to ensure both strong reasoning and efficient decision-making capabilities via unified tokenization of multimodal interaction data. First, we introduce a self-supervised approach to learn a behavior encoder that produces discretized tokens for behavior trajectories $\tau = \{o_0, a_0, \dots\}$ and an imitation learning policy decoder conditioned on these tokens. These additional behavior tokens will be augmented to the vocabulary of pretrained Multimodal Language Models. With this encoder, we then pack long-term multimodal interactions involving task instructions, memories, thoughts, observations, textual responses, behavior trajectories, etc into unified token sequences and model them with autoregressive transformers. Thanks to the semantically meaningful behavior tokens, the resulting VLA model, OmniJARVIS, can reason (by producing chain-of-thoughts), plan, answer questions, and act (by producing behavior tokens for the imitation learning policy decoder). OmniJARVIS demonstrates excellent performances on a comprehensive collection of atomic, programmatic, and open-ended tasks in open-world Minecraft. Our analysis further unveils the crucial design principles in interaction data formation, unified tokenization, and its scaling potentials. The dataset, models, and code will be released at https://craftjarvis.org/OmniJARVIS.
Shaofei Cai, Zhancun Mu, Haowei Lin, Ceyao Zhang, Xuejie Liu, Qing Li 0003, Anji Liu, Xiaojian Ma 0001, Yitao Liang
NeurIPS4
2024 TFG: Unified Training-Free Guidance for Diffusion Models
abstract
Given an unconditional diffusion model and a predictor for a target property of interest (e.g., a classifier), the goal of training-free guidance is to generate samples with desirable target properties without additional training. Existing methods, though effective in various individual applications, often lack theoretical grounding and rigorous testing on extensive benchmarks. As a result, they could even fail on simple tasks, and applying them to a new problem becomes unavoidably difficult. This paper introduces a novel algorithmic framework encompassing existing methods as special cases, unifying the study of training-free guidance into the analysis of an algorithm-agnostic design space. Via theoretical and empirical investigation, we propose an efficient and effective hyper-parameter searching strategy that can be readily applied to any downstream task. We systematically benchmark across 7 diffusion models on 16 tasks with 40 targets, and improve performance by 8.5% on average. Our framework and benchmark offer a solid foundation for conditional generation in a training-free manner.
Haotian Ye, Haowei Lin, Jiaqi Han 0001, Minkai Xu, Yitao Liang, Jianzhu Ma, James Zou 0001, Stefano Ermon
NeurIPS2
2023 FLatS: Principled Out-of-Distribution Detection with Feature-Based Likelihood Ratio Score
abstract
Detecting out-of-distribution (OOD) instances is crucial for NLP models in practical applications.Although numerous OOD detection methods exist, most of them are empirical.Backed by theoretical analysis, this paper advocates for the measurement of the "OOD-ness" of a test case x through the likelihood ratio between out-distribution P out and in-distribution P in .We argue that the state-of-the-art (SOTA) feature-based OOD detection methods, such as Maha (Lee et al., 2018) and KNN (Sun et al., 2022), are suboptimal since they only estimate in-distribution density p in (x).To address this issue, we propose FLatS, a principled solution for OOD detection based on likelihood ratio.Moreover, we demonstrate that FLatS can serve as a general framework capable of enhancing other OOD detection methods by incorporating out-distribution density p out (x) estimation.Experiments show that FLatS establishes a new SOTA on popular benchmarks. 1
Haowei Lin, Yuntian Gu
EMNLP1
2023 Continual Pre-training of Language Models
Zixuan Ke, Yijia Shao, Haowei Lin, Tatsuya Konishi, Gyuhak Kim, Bing Liu 0001
ICLR3
2022 Continual Training of Language Models for Few-Shot Learning
abstract
Recent work on applying large language models (LMs) achieves impressive performance in many NLP applications.Adapting or posttraining an LM using an unlabeled domain corpus can produce even better performance for end-tasks in the domain.This paper proposes the problem of continually extending an LM by incrementally post-train the LM with a sequence of unlabeled domain corpora to expand its knowledge without forgetting its previous skills.The goal is to improve the few-shot end-task learning in these domains.The resulting system is called CPT (Continual Post-Training), which to our knowledge, is the first continual post-training system.Experimental results verify its effectiveness.
Zixuan Ke, Haowei Lin, Yijia Shao, Hu Xu 0001, Lei Shu 0004, Bing Liu 0001
EMNLP2
2022 Adapting a Language Model While Preserving its General Knowledge
abstract
Domain-adaptive pre-training (or DA-training for short), also known as post-training, aims to train a pre-trained general-purpose language model (LM) using an unlabeled corpus of a particular domain to adapt the LM so that endtasks in the domain can give improved performances.However, existing DA-training methods are in some sense blind as they do not explicitly identify what knowledge in the LM should be preserved and what should be changed by the domain corpus.This paper shows that the existing methods are suboptimal and proposes a novel method to perform a more informed adaptation of the knowledge in the LM by (1) soft-masking the attention heads based on their importance to best preserve the general knowledge in the LM and (2) contrasting the representations of the general and the full (both general and domain knowledge) to learn an integrated representation with both general and domain-specific knowledge.Experimental results will demonstrate the effectiveness of the proposed approach.1
Zixuan Ke, Yijia Shao, Haowei Lin, Hu Xu 0001, Lei Shu 0004, Bing Liu 0001
EMNLP3
2022 CMG: A Class-Mixed Generation Approach to Out-of-Distribution Detection
Mengyu Wang 0002, Yijia Shao, Haowei Lin, Wenpeng Hu, Bing Liu 0001
ECML/PKDD (4)3
2021 Particle swarm optimized neural networks based local tracking control scheme of unknown nonlinear interconnected systems
Bo Zhao 0015, Fangchao Luo, Haowei Lin, Derong Liu 0001
Neural Networks3
2019 Local Near-Optimal Control for Interconnected Systems with Time-Varying Delays
Qiuye Wu, Haowei Lin, Bo Zhao 0015, Derong Liu 0001
ICONIP (2)2
2019 Energy Consumption of IT System in Cloud Data Center: Architecture, Factors and Prediction
Haowei Lin, Xiaolong Xu 0002, Xinheng Wang 0001
NPC1