Haiteng Zhao

dblp:304/8330 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
11since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 4 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 MUR: Momentum Uncertainty guided Reasoning for Large Language Models
abstract
Hang Yan, Fangzhi Xu, Rongman Xu, Yifei Li, Jian Zhang, Haoran Luo, Xiaobao Wu, Anh Tuan Luu, Haiteng Zhao, Qika Lin, Jun Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Hang Yan 0010, Fangzhi Xu, Rongman Xu, Yifei Li 0006, Jian Zhang 0087, Haoran Luo 0001, Xiaobao Wu, Anh Tuan Luu, Haiteng Zhao, Qika Lin, Jun Liu 0002
ACL (1)9
2025 φ-Decoding: Adaptive Foresight Sampling for Balanced Inference-Time Exploration and Exploitation
Fangzhi Xu, Hang Yan 0010, Haiteng Zhao, Jun Liu 0002, Qika Lin, Zhiyong Wu 0003
ACL (1)4
2025 Genius: A Generalizable and Purely Unsupervised Self-Training Framework For Advanced Reasoning
abstract
Fangzhi Xu, Hang Yan, Chang Ma, Haiteng Zhao, Qiushi Sun, Kanzhi Cheng, Junxian He, Jun Liu, Zhiyong Wu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Fangzhi Xu, Hang Yan 0010, Haiteng Zhao, Qiushi Sun, Kanzhi Cheng, Junxian He, Jun Liu 0002, Zhiyong Wu 0003
ACL (1)4
2025 Instruction-Based Molecular Graph Generation with Unified Text-Graph Diffusion Model
abstract
Recent advancements in computational chemistry have increasingly emphasized generating and editing molecules from textual instructions. However, integrating graph generation with instruction understanding remains challenging, as most existing approaches either rely on molecular sequences in text modality with limited structural information, or struggle with multimodal alignment in graph diffusion methods. To address these limitations, we propose UTGDiff (Unified Text-Graph Diffusion Model), a novel framework that utilizes pre-trained language models for discrete graph diffusion, enabling the generation of molecular graphs from instructions. UTGDiff introduces a unified text-graph transformer as a denoising network, adapted with minimal modifications from language models to process graph data via attention bias. Experimental results show that UTGDiff consistently outperforms both sequence-based and conditional graph-diffusion baselines on instruction-based molecule generation and editing tasks with fewer parameters, covering instructions specifying molecular structures or properties.
Yuran Xiang, Haiteng Zhao, Zhi-Hong Deng 0001
ECAI2
2025 Non-myopic Generation of Language Models for Reasoning and Planning
abstract
Large Language Models (LLMs) have demonstrated remarkable abilities in reasoning and planning. Despite their success in various domains, such as mathematical problem-solving and coding, LLMs face challenges in ensuring reliable and optimal planning due to the inherent myopic nature of autoregressive decoding. This paper revisits LLM reasoning from an optimal control perspective, proposing a novel method, Predictive-Decoding, that leverages Model Predictive Control to enhance planning accuracy. By reweighting LLM distributions based on foresight trajectories, Predictive-Decoding aims to mitigate early errors and promote non-myopic planning. Our experiments show significant improvements across a wide range of tasks in math, coding, and agent-based scenarios. Furthermore, Predictive-Decoding demonstrates computational efficiency, outperforming search baselines while utilizing inference compute more effectively. This study provides insights into optimizing LLM planning capabilities.
Haiteng Zhao, Junlei Zhang, Junxian He, Lingpeng Kong
ICLR2
2025 Semi-off-Policy Reinforcement Learning for Vision-Language Slow-Thinking Reasoning
abstract
Enhancing large vision-language models (LVLMs) with visual slow-thinking reasoning is crucial for solving complex multimodal tasks. However, since LVLMs are mainly trained with vision-language alignment, it is difficult to adopt on-policy reinforcement learning (RL) to develop the slow thinking ability because the rollout space is restricted by its initial abilities. Off-policy RL offers a way to go beyond the current policy, but directly distilling trajectories from external models may cause visual hallucinations due to mismatched visual perception abilities across models. To address these issues, this paper proposes **SOPHIA**, a simple and scalable **S**emi-**O**ff-**P**olicy RL for vision-language slow-t**HI**nking re**A**soning. SOPHIA builds a semi-off-policy behavior model by combining on-policy visual understanding from a trainable LVLM with off-policy slow-thinking reasoning from a language model, assigns outcome-based rewards to reasoning, and propagates visual rewards backward. Then LVLM learns slow-thinking reasoning ability from the obtained reasoning trajectories using propagated rewards via off-policy RL algorithms. Extensive experiments with InternVL2.5 and InternVL3.0 with 8B and 38B sizes show the effectiveness of SOPHIA. Notably, SOPHIA improves InternVL3.0-38B by 8.50\% in average, reaching state-of-the-art performance among open-source LVLMs on multiple multimodal reasoning benchmarks, and even outperforms some closed-source models (e.g., GPT-4.1) on the challenging MathVision and OlympiadBench, achieving 49.08\% and 49.95\% pass@1 accuracy, respectively. Analysis shows SOPHIA outperforms supervised fine-tuning and direct on-policy RL methods, offering a better policy initialization for further on-policy training.
Haiteng Zhao, Yuzhe Gu, Songyang Gao, Kuikun Liu, Haian Huang, Jianfei Gao 0003, Dahua Lin, Kai Chen 0026
NeurIPS2
2024 Retrieved Sequence Augmentation for Protein Representation Learning
abstract
Chang Ma, Haiteng Zhao, Lin Zheng, Jiayi Xin, Qintong Li, Lijun Wu, Zhihong Deng, Yang Young Lu, Qi Liu, Sheng Wang, Lingpeng Kong. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Haiteng Zhao, Jiayi Xin, Qintong Li, Zhi-Hong Deng 0001, Qi Liu 0049, Lingpeng Kong
EMNLP2
2023 Are More Layers Beneficial to Graph Transformers?
Haiteng Zhao, Shuming Ma, Dongdong Zhang 0001, Zhi-Hong Deng 0001, Furu Wei
ICLR1
2023 GIMLET: A Unified Graph-Text Model for Instruction-Based Molecule Zero-Shot Learning
abstract
Molecule property prediction has gained significant attention in recent years. The main bottleneck is the label insufficiency caused by expensive lab experiments. In order to alleviate this issue and to better leverage textual knowledge for tasks, this study investigates the feasibility of employing natural language instructions to accomplish molecule-related tasks in a zero-shot setting. We discover that existing molecule-text models perform poorly in this setting due to inadequate treatment of instructions and limited capacity for graphs. To overcome these issues, we propose GIMLET, which unifies language models for both graph and text data. By adopting generalized position embedding, our model is extended to encode both graph structures and instruction text without additional graph encoding modules. GIMLET also decouples encoding of the graph from tasks instructions in the attention mechanism, enhancing the generalization of graph features across novel tasks. We construct a dataset consisting of more than two thousand molecule tasks with corresponding instructions derived from task descriptions. We pretrain GIMLET on the molecule tasks along with instructions, enabling the model to transfer effectively to a broad range of tasks. Experimental results demonstrate that GIMLET significantly outperforms molecule-text baselines in instruction-based zero-shot learning, even achieving closed results to supervised GNN models on tasks such as toxcast and muv.
Haiteng Zhao, Shengchao Liu, Hannan Xu, Jie Fu 0001, Zhi-Hong Deng 0001, Lingpeng Kong, Qi Liu 0049
NeurIPS1
2022 Certified Robustness Against Natural Language Attacks by Causal Intervention
abstract
Deep learning models have achieved great success in many fields, yet they are vulnerable to adversarial examples. This paper follows a causal perspective to look into the adversarial vulnerability and proposes Causal Intervention by Semantic Smoothing (CISS), a novel framework towards robustness against natural language attacks. Instead of merely fitting observational data, CISS learns causal effects p(y|do(x)) by smoothing in the latent semantic space to make robust predictions, which scales to deep architectures and avoids tedious construction of noise customized for specific attacks. CISS is provably robust against word substitution attacks, as well as empirically robust even when perturbations are strengthened by unknown attack algorithms. For example, on YELP, CISS surpasses the runner-up by 6.8% in terms of certified robustness against word substitutions, and achieves 80.7% empirical robustness when syntactic attacks are integrated.
Haiteng Zhao, Xinshuai Dong, Anh Tuan Luu, Zhi-Hong Deng 0001, Hanwang Zhang
ICML1
2022 Domain Adaptation via Maximizing Surrogate Mutual Information
abstract
Unsupervised domain adaptation (UDA), which is an important topic in transfer learning, aims to predict unlabeled data from target domain with access to labeled data from the source domain. In this work, we propose a novel framework called SIDA (Surrogate Mutual Information Maximization Domain Adaptation) with strong theoretical guarantees. To be specific, SIDA implements adaptation by maximizing mutual information (MI) between features. In the framework, a surrogate joint distribution models the underlying joint distribution of the unlabeled target domain. Our theoretical analysis validates SIDA by bounding the expected risk on target domain with MI and surrogate distribution bias. Experiments show that our approach is comparable with state-of-the-art unsupervised adaptation methods on standard UDA tasks.
Haiteng Zhao, Qinyu Chen, Zhi-Hong Deng 0001
IJCAI1