Haizhen Huang

dblp:304/7795 · DBLP profile ↗
← Back
14ranked-venue papers
1as first author
14since 2021 · last 2026
0009-0005-7145-2500ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 12 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Too Long, Do Re-weighting for Efficient LLM Reasoning Compression
abstract
Zhong-Zhi Li, Xiao Liang, Zihao Tang, Lei Ji, Peijie Wang, Haotian Xu, Xing W, Haizhen Huang, Weiwei Deng, Yeyun Gong, Zhijiang Guo, Xiao Liu, Fei Yin, Cheng-Lin Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zhongzhi Li, Lei Ji 0001, Xing W, Haizhen Huang, Yeyun Gong, Zhijiang Guo, Xiao Liu 0029, Cheng-Lin Liu 0001
ACL (1)8
2026 Mnemis: Dual-Route Retrieval on Hierarchical Graphs for Long-Term LLM Memory
abstract
Zihao Tang, Xin Yu, Ziyu Xiao, Zengxuan Wen, Zelin Li, Jiaxi Zhou, Hualei Wang, Haohua Wang, Haizhen Huang, Weiwei Deng, Feng Sun, Qi Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Ziyu Xiao, Zengxuan Wen, Hualei Wang, Haizhen Huang, Feng Sun 0008, Qi Zhang 0066
ACL (1)9
2025 NL2Lean: Translating Natural Language into Lean 4 through Multi-Aspect Reinforcement Learning
abstract
Yue Fang, Shaohan Huang, Xin Yu, Haizhen Huang, Zihan Zhang, Weiwei Deng, Furu Wei, Feng Sun, Qi Zhang, Zhi Jin. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Yue Fang 0001, Shaohan Huang, Haizhen Huang, Furu Wei, Feng Sun 0008, Qi Zhang 0066, Zhi Jin 0001
EMNLP4
2025 VirB: A Virus Hierarchical Classification Method Based on ModernBERT
Haizhen Huang, Haodi Feng, Daming Zhu
ICIC (26)1
2024 Text Diffusion with Reinforced Conditioning
abstract
Diffusion models have demonstrated exceptional capability in generating high-quality images, videos, and audio. Due to their adaptiveness in iterative refinement, they provide a strong potential for achieving better non-autoregressive sequence generation. However, existing text diffusion models still fall short in their performance due to a challenge in handling the discreteness of language. This paper thoroughly analyzes text diffusion models and uncovers two significant limitations: degradation of self-conditioning during training and misalignment between training and sampling. Motivated by our findings, we propose a novel Text Diffusion model called TReC, which mitigates the degradation with Reinforced Conditioning and the misalignment by Time-Aware Variance Scaling. Our extensive experiments demonstrate the competitiveness of TReC against autoregressive, non-autoregressive, and diffusion baselines. Moreover, qualitative analysis shows its advanced ability to fully utilize the diffusion process in refining samples.
Yuxuan Liu 0011, Tianchi Yang, Shaohan Huang, Haizhen Huang, Furu Wei, Feng Sun 0008, Qi Zhang 0066
AAAI5
2024 HD-Eval: Aligning Large Language Model Evaluators Through Hierarchical Criteria Decomposition
abstract
Yuxuan Liu, Tianchi Yang, Shaohan Huang, Zihan Zhang, Haizhen Huang, Furu Wei, Weiwei Deng, Feng Sun, Qi Zhang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Yuxuan Liu 0011, Tianchi Yang, Shaohan Huang, Haizhen Huang, Furu Wei, Feng Sun 0008, Qi Zhang 0066
ACL (1)5
2024 Calibrating LLM-Based Evaluator
abstract
Recent advancements in large language models (LLMs) and their emergent capabilities make LLM a promising reference-free evaluator on the quality of natural language generation, and a competent alternative to human evaluation. However, hindered by the closed-source or high computational demand to host and tune, there is a lack of practice to further calibrate an off-the-shelf LLM-based evaluator towards better human alignment. In this work, we propose AutoCalibrate, a multi-stage, gradient-free approach to automatically calibrate and align an LLM-based evaluator toward human preference. Instead of explicitly modeling human preferences, we first implicitly encompass them within a set of human labels. Then, an initial set of scoring criteria is drafted by the language model itself, leveraging in-context learning on different few-shot examples. To further calibrate this set of criteria, we select the best performers and re-draft them with self-refinement. Our experiments on multiple text quality evaluation datasets illustrate a significant improvement in correlation with expert evaluation through calibration. Our comprehensive qualitative analysis conveys insightful intuitions and observations on the essence of effective scoring criteria.
Yuxuan Liu 0011, Tianchi Yang, Shaohan Huang, Haizhen Huang, Furu Wei, Feng Sun 0008, Qi Zhang 0066
LREC/COLING5
2023 Dual-Alignment Pre-training for Cross-lingual Sentence Embedding
abstract
Ziheng Li, Shaohan Huang, Zihan Zhang, Zhi-Hong Deng, Qiang Lou, Haizhen Huang, Jian Jiao, Furu Wei, Weiwei Deng, Qi Zhang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Ziheng Li 0003, Shaohan Huang, Zhi-Hong Deng 0001, Qiang Lou, Haizhen Huang, Jian Jiao 0007, Furu Wei, Qi Zhang 0066
ACL (1)6
2023 Towards Better Entity Linking with Multi-View Enhanced Distillation
abstract
Yi Liu, Yuan Tian, Jianxun Lian, Xinlong Wang, Yanan Cao, Fang Fang, Wen Zhang, Haizhen Huang, Weiwei Deng, Qi Zhang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Yi Liu 0067, Jianxun Lian, Yanan Cao 0001, Fang Fang 0009, Haizhen Huang, Qi Zhang 0066
ACL (1)8
2023 Democratizing Reasoning Ability: Tailored Learning from Large Language Model
abstract
Zhaoyang Wang, Shaohan Huang, Yuxuan Liu, Jiahai Wang, Minghui Song, Zihan Zhang, Haizhen Huang, Furu Wei, Weiwei Deng, Feng Sun, Qi Zhang. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Shaohan Huang, Yuxuan Liu 0011, Jiahai Wang, Minghui Song, Haizhen Huang, Furu Wei, Feng Sun 0008, Qi Zhang 0066
EMNLP7
2023 PASS: Personalized Advertiser-aware Sponsored Search
abstract
The nucleus of online sponsored search systems lies in measuring the relevance between the search intents of users and the advertising purposes of advertisers. Existing conventional doublet-based (query-keyword) relevance models solely rely on short queries and keywords to uncover such intents, which ignore the diverse and personalized preferences of participants (i.e., users and advertisers), resulting in undesirable advertising performance. In this paper, we investigate the novel problem of Personalized A dvertiser-aware Sponsored Search (PASS). Our motivation lies in incorporating the portraits of users and advertisers into relevance models to facilitate the modeling of intrinsic search intents and advertising purposes, leading to a quadruple-based (i.e., user-query-keyword-advertiser) task. Various types of historical behaviors are explored in the format of hypergraphs to provide abundant signals on identifying the preferences of participants. A novel heterogeneous textual hypergraph transformer is further proposed to deeply fuse the textual semantics and the high-order hypergraph topology. Our proposal is extensively evaluated over real industry datasets, and experimental results demonstrate its superiority.
Zhoujin Tian, Chaozhuo Li, Zhiqiang Zuo 0004, Zengxuan Wen, Lichao Sun 0001, Xinyue Hu 0003, Haizhen Huang, Senzhang Wang, Xing Xie 0001, Qi Zhang 0066
KDD8
2023 Multi-Grained Topological Pre-Training of Language Models in Sponsored Search
abstract
Relevance models measure the semantic closeness between queries and the candidate ads, widely recognized as the nucleus of sponsored search systems. Conventional relevance models solely rely on the textual data within the queries and ads, whose performance is hindered by the scarce semantic information in these short texts. Recently, user behavior graphs have been incorporated to provide complementary information beyond pure textual semantics.Despite the promising performance, behavior-enhanced models suffer from exhausting resource costs due to the extra computations introduced by explicit topological aggregations. In this paper, we propose a novel Multi-Grained Topological Pre-Training paradigm, MGTLM, to teach language models to understand multi-grained topological information in behavior graphs, which contributes to eliminating explicit graph aggregations and avoiding information loss. Extensive experimental results over online and offline settings demonstrate the superiority of our proposal.
Zhoujin Tian, Chaozhuo Li, Zhiqiang Zuo 0004, Zengxuan Wen, Xinyue Hu 0003, Haizhen Huang, Senzhang Wang, Xing Xie 0001, Qi Zhang 0066
SIGIR7
2022 PromptBERT: Improving BERT Sentence Embeddings with Prompts
abstract
Ting Jiang, Jian Jiao, Shaohan Huang, Zihan Zhang, Deqing Wang, Fuzhen Zhuang, Furu Wei, Haizhen Huang, Denvy Deng, Qi Zhang. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Jian Jiao 0007, Shaohan Huang, Fuzhen Zhuang, Furu Wei, Haizhen Huang, Denvy Deng, Qi Zhang 0066
EMNLP8
2022 RAPO: An Adaptive Ranking Paradigm for Bilingual Lexicon Induction
abstract
Zhoujin Tian, Chaozhuo Li, Shuo Ren, Zhiqiang Zuo, Zengxuan Wen, Xinyue Hu, Xiao Han, Haizhen Huang, Denvy Deng, Qi Zhang, Xing Xie. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Zhoujin Tian, Chaozhuo Li, Shuo Ren 0002, Zhiqiang Zuo 0004, Zengxuan Wen, Xinyue Hu 0003, Haizhen Huang, Denvy Deng, Qi Zhang 0066, Xing Xie 0001
EMNLP8