Haozhe Zhao

dblp:299/7199 · DBLP profile ↗
← Back
16ranked-venue papers
5as first author
16since 2021 · last 2026
0000-0003-0502-4426ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 5 first-author · 15 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021
YearPublicationVenuePosition
2026 Teaching Large Language Models to Maintain Contextual Faithfulness via Synthetic Tasks and Reinforcement Learning
abstract
Teaching large language models (LLMs) to be faithful in the provided context is crucial for building reliable information-seeking systems. Therefore, we propose a systematic framework, CANOE, to reduce faithfulness hallucinations of LLMs across different downstream tasks without human annotations. Specifically, we first synthesize short-form question-answering (QA) data with four diverse tasks to construct high-quality and easily verifiable training data without human annotation. Also, we propose Dual-GRPO, a rule-based reinforcement learning method that includes three tailored rule-based rewards derived from synthesized short-form QA data, while simultaneously optimizing both short-form and long-form response generation. Notably, Dual-GRPO eliminates the need to manually label preference data to train reward models and avoids over-optimizing short-form generation when relying only on the synthesized short-form QA data. Experimental results show that CANOE greatly improves the faithfulness of LLMs across 11 different tasks, even outperforming the most advanced LLMs, e.g., GPT-4o and OpenAI o1.
Shuzheng Si, Haozhe Zhao, Yuzhuo Bai, Zhitong Wang, Bofei Gao, Kangyang Luo, Wenhao Li 0003, Yufei Huang 0008, Gang Chen 0039, Fanchao Qi, Minjia Zhang, Baobao Chang, Maosong Sun 0001
AAAI2
2026 A Goal Without a Plan Is Just a Wish: Efficient and Effective Global Planner Training for Long-Horizon Agent Task
abstract
Shuzheng Si, Haozhe Zhao, Kangyang Luo, Gang Chen, Fanchao Qi, Minjia Zhang, Baobao Chang, Maosong Sun. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Shuzheng Si, Haozhe Zhao, Kangyang Luo, Gang Chen 0039, Fanchao Qi, Minjia Zhang, Baobao Chang, Maosong Sun 0001
ACL (1)2
2026 SADGCN-GC: self-attention-based deep graph convolutional neural network with quantization and visualization for graph classification
Zhaofeng Chen, Baodan Ye, Junhan She, Haozhe Zhao, Naixuan Guo, Binghe Sun
Vis. Comput.5
2025 Aligning Large Language Models to Follow Instructions and Hallucinate Less via Effective Data Filtering
abstract
Shuzheng Si, Haozhe Zhao, Gang Chen, Cheng Gao, Yuzhuo Bai, Zhitong Wang, Kaikai An, Kangyang Luo, Chen Qian, Fanchao Qi, Baobao Chang, Maosong Sun. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Shuzheng Si, Haozhe Zhao, Gang Chen 0039, Yuzhuo Bai, Zhitong Wang, Kaikai An, Kangyang Luo, Fanchao Qi, Baobao Chang, Maosong Sun 0001
ACL (1)2
2025 CCAgent: Coordinating Collaborative Data Scaling for Operating System Agents via Web3
Liang Chen 0024, Haozhe Zhao, Yinzhen Huang, Tsekai Lin, Weichu Xie, Peiyi Wang, Runxin Xu, Ming Wu 0007, Baobao Chang
CIKM2
2025 GATEAU: Selecting Influential Samples for Long Context Alignment
abstract
Shuzheng Si, Haozhe Zhao, Gang Chen, Yunshui Li, Kangyang Luo, Chuancheng Lv, Kaikai An, Fanchao Qi, Baobao Chang, Maosong Sun. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Shuzheng Si, Haozhe Zhao, Gang Chen 0039, Yunshui Li, Kangyang Luo, Chuancheng Lv, Kaikai An, Fanchao Qi, Baobao Chang, Maosong Sun 0001
EMNLP2
2025 Looking Beyond Text: Reducing Language Bias in Large Vision-Language Models via Multimodal Dual-Attention and Soft-Image Guidance
abstract
Large vision-language models (LVLMs) have achieved impressive results in vision-language tasks.However, LVLMs suffer from hallucinations caused by language bias, which neglects images while over-relying on text.We identify two reasons for the bias: 1).Different training scales between the LLM pretraining and LVLM alignment stage.2).The learned inference bias due to short-term dependency of text data.Therefore, we propose LACING, designed to address such bias with MuLtimodal DuAlattention MeChanIsm (MDA) aNd Soft-Image Guidance (SIG).Specifically, MDA adopts a parallel dual-attention mechanism that constructs separate attention for visual and text inputs to enhance integration of visual inputs across model.SIG uses a learnable soft visual prompt during training and inference to replace visual inputs, designed to compel LVLMs to prioritize text inputs during inference.Experiments across different model architectures and scales demonstrate that LACING effectively debiases LVLMs from their language bias, enhancing visual comprehension and reducing hallucinations without additional resources.
Haozhe Zhao, Shuzheng Si, Liang Chen 0024, Yichi Zhang 0010, Maosong Sun 0001, Baobao Chang, Minjia Zhang
EMNLP1
2025 A Spark of Vision-Language Intelligence: 2-Dimensional Autoregressive Transformer for Efficient Finegrained Image Generation
abstract
This work tackles the information loss bottleneck of vector-quantization (VQ) autoregressive image generation by introducing a novel model architecture called the 2-Dimensional Autoregression (DnD) Transformer. The DnD-Transformer predicts more codes for an image by introducing a new direction, **model depth**, along with the sequence length. Compared to 1D autoregression and previous work using similar 2D image decomposition such as RQ-Transformer, the DnD-Transformer is an end-to-end model that can generate higher quality images with the same backbone model size and sequence length, opening a new optimization perspective for autoregressive image generation. Furthermore, our experiments reveal that the DnD-Transformer's potential extends beyond generating natural images. It can even generate images with rich text and graphical elements in a self-supervised manner, demonstrating an understanding of these combined modalities. This has not been previously demonstrated for popular vision generative models such as diffusion models, showing a spark of vision-language intelligence when trained solely on images. Code, datasets and models are open at https://github.com/chenllliang/DnD-Transformer.
Liang Chen 0024, Sinan Tan, Zefan Cai, Weichu Xie, Haozhe Zhao, Yichi Zhang 0010, Junyang Lin, Jinze Bai, Tianyu Liu 0001, Baobao Chang
ICLR5
2025 MMEvalPro: Calibrating Multimodal Benchmarks Towards Trustworthy and Efficient Evaluation
abstract
Jinsheng Huang, Liang Chen, Taian Guo, Fu Zeng, Yusheng Zhao, Bohan Wu, Ye Yuan, Haozhe Zhao, Zhihui Guo, Yichi Zhang, Jingyang Yuan, Wei Ju, Luchen Liu, Tianyu Liu, Baobao Chang, Ming Zhang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Jinsheng Huang, Liang Chen 0024, Taian Guo, Fu Zeng, Yusheng Zhao, Bohan Wu, Ye Yuan 0016, Haozhe Zhao, Zhihui Guo, Yichi Zhang 0010, Jingyang Yuan, Wei Ju 0001, Luchen Liu, Tianyu Liu 0001, Baobao Chang, Ming Zhang 0004
NAACL (Long Papers)8
2025 NEP: Autoregressive Image Editing via Next Editing Token Prediction
abstract
Text-guided image editing involves modifying a source image based on a language instruction and, typically, requires changes to only small local regions. However, existing approaches generate the entire target image rather than selectively regenerate only the intended editing areas. This results in (1) unnecessary computational costs and (2) a bias toward reconstructing non-editing regions, which compromises the quality of the intended edits. To resolve these limitations, we propose to formulate image editing as $\textbf{N}$ext $\textbf{E}$diting-token $\textbf{P}$rediction (NEP) based on autoregressive image generation, where only regions that need to be edited are regenerated, thus avoiding unintended modification to the non-editing areas. To enable any-region editing, we propose to pre-train an any-order autoregressive text-to-image (T2I) model. Once trained, it is capable of zero-shot image editing and can be easily adapted to NEP for image editing, which achieves a new state-of-the-art on widely used image editing benchmarks. Moreover, our model naturally supports test-time scaling (TTS) through iteratively refining its generation in a zero-shot manner.
Huimin Wu 0001, Xiaojian Ma 0001, Haozhe Zhao, Yanpeng Zhao, Qing Li 0003
NeurIPS3
2024 An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
Liang Chen 0024, Haozhe Zhao, Tianyu Liu 0001, Shuai Bai, Junyang Lin, Chang Zhou 0005, Baobao Chang
ECCV (81)2
2024 MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning
abstract
Since the resurgence of deep learning, vision-language models (VLMs) enhanced by large language models (LLMs) have grown exponentially in popularity. However, while LLMs can utilize extensive background knowledge and task information with in-context learning, most VLMs still struggle with understanding complex multi-modal prompts with multiple images, making VLMs less effective in downstream vision-language tasks. In this paper, we address the limitation above by 1) introducing vision-language Model with **M**ulti-**M**odal **I**n-**C**ontext **L**earning(MMICL), a new approach to allow the VLM to deal with multi-modal inputs efficiently; 2) proposing a novel context scheme to augment the in-context learning ability of the VLM; 3) constructing the Multi-modal In-Context Learning (MIC) dataset, designed to enhance the VLM's ability to understand complex multi-modal prompts. Our experiments confirm that MMICL achieves new state-of-the-art zero-shot performance on a wide range of general vision-language tasks, especially for complex benchmarks, including MME and MMBench. Our analysis demonstrates that MMICL effectively tackles the challenge of complex multi-modal prompt understanding and emerges the impressive ICL ability. Furthermore, we observe that MMICL successfully alleviates language bias in VLMs, a common issue for VLMs that often leads to hallucination when faced with extensive textual context. Our code, dataset, dataset tool, and model are available at https://github.com/PKUnlp-icler/MIC.
Haozhe Zhao, Zefan Cai, Shuzheng Si, Xiaojian Ma 0001, Kaikai An, Liang Chen 0024, Zixuan Liu 0001, Sheng Wang 0012, Wenjuan Han, Baobao Chang
ICLR1
2024 Mitigating Language-Level Performance Disparity in mPLMs via Teacher Language Selection and Cross-lingual Self-Distillation
abstract
Haozhe Zhao, Zefan Cai, Shuzheng Si, Liang Chen, Yufeng He, Kaikai An, Baobao Chang. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Haozhe Zhao, Zefan Cai, Shuzheng Si, Liang Chen 0024, Yufeng He, Kaikai An, Baobao Chang
NAACL-HLT1
2024 UltraEdit: Instruction-based Fine-Grained Image Editing at Scale
abstract
This paper presents UltraEdit, a large-scale (~ 4M editing samples), automatically generated dataset for instruction-based image editing. Our key idea is to address the drawbacks in existing image editing datasets like InstructPix2Pix and MagicBrush, and provide a systematic approach to producing massive and high-quality image editing samples: 1) UltraEdit includes more diverse editing instructions by combining LLM creativity and in-context editing examples by human raters; 2) UltraEdit is anchored on real images (photographs or artworks), which offers more diversity and less biases than those purely synthesized by text-to-image models; 3) UltraEdit supports region-based editing with high-quality, automatically produced region annotations. Our experiments show that canonical diffusion-based editing baselines trained on UltraEdit set new records on challenging MagicBrush and Emu-Edit benchmarks, respectively. Our analysis further confirms the crucial role of real image anchors and region-based editing data. The dataset, code, and models will be made public.
Haozhe Zhao, Xiaojian Ma 0001, Liang Chen 0024, Shuzheng Si, Rujie Wu, Kaikai An, Peiyu Yu, Minjia Zhang, Qing Li 0003, Baobao Chang
NeurIPS1
2023 Removing Camouflage and Revealing Collusion: Leveraging Gang-crime Pattern in Fraudster Detection
abstract
As one of the major threats to the healthy development of various online platforms, fraud has become increasingly committed in the form of gangs since collusive fraudulent activities are much easier to obtain illicit benefits with lower exposure risk. To detect fraudsters in a gang, spatio-temporal graph neural network models have been widely applied to detect both temporal and spatial collusive patterns. However, a closer peek into real-world records of fraudsters can reveal that fraud gangs usually conduct community-level camouflage, specified by two types, i.e., temporal and spatial camouflage. Such camouflage can disguise gangs as benign communities by concealing collusive patterns and thus deceiving many existing graph neural network models. In the meantime, many existing graph neural network models suffer from the challenge of extreme sample imbalance caused by rare fraudsters hidden among massive users. To handle all these challenges, in this paper, we propose a generative adversarial network framework, named Adversarial Camouflage Detector, to detect fraudsters. Concretely, this ACD framework consists of four modules, in charge of community division, camouflage identification, fraudster detection, and camouflage generation, respectively. The first three modules form up a discriminator that uses spatio-temporal graph neural networks as the foundation model and enhance fraudster detection by amplifying the gangs' collusive patterns through automatically identifying and removing camouflage. Meanwhile, the camouflage generation module plays as the generator role that generates fraudsters samples by competing against the discriminator to alleviate the challenge of sample imbalance and increase the model robustness. The experimental result shows that our proposed method outperforms other methods on real-world datasets.
Lewen Wang, Haozhe Zhao, Cunguang Feng, Weiqing Liu, Congrui Huang, Marco Santoni, Manuel Cristofaro, Paola Jafrancesco, Jiang Bian 0002
KDD2
2021 Traffic Accident Prediction Methods Based on Multi-factor Models
Haozhe Zhao, Guozheng Rao
KSEM1