VLDB 2026 Research / reviewers in the wild / expert
Haizhen Huang
dblp:304/7795
· DBLP profile ↗
14ranked-venue papers
1as first author
14since 2021 · last 2026
0009-0005-7145-2500ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 12 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Too Long, Do Re-weighting for Efficient LLM Reasoning CompressionabstractZhong-Zhi Li, Xiao Liang, Zihao Tang, Lei Ji, Peijie Wang, Haotian Xu, Xing W, Haizhen Huang, Weiwei Deng, Yeyun Gong, Zhijiang Guo, Xiao Liu, Fei Yin, Cheng-Lin Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zhongzhi Li, Lei Ji 0001, Xing W, Haizhen Huang, Yeyun Gong, Zhijiang Guo, Xiao Liu 0029, Cheng-Lin Liu 0001 |
ACL (1) | 8 |
| 2026 | Mnemis: Dual-Route Retrieval on Hierarchical Graphs for Long-Term LLM MemoryabstractZihao Tang, Xin Yu, Ziyu Xiao, Zengxuan Wen, Zelin Li, Jiaxi Zhou, Hualei Wang, Haohua Wang, Haizhen Huang, Weiwei Deng, Feng Sun, Qi Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Ziyu Xiao, Zengxuan Wen, Hualei Wang, Haizhen Huang, Feng Sun 0008, Qi Zhang 0066 |
ACL (1) | 9 |
| 2025 | NL2Lean: Translating Natural Language into Lean 4 through Multi-Aspect Reinforcement LearningabstractYue Fang, Shaohan Huang, Xin Yu, Haizhen Huang, Zihan Zhang, Weiwei Deng, Furu Wei, Feng Sun, Qi Zhang, Zhi Jin. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Yue Fang 0001, Shaohan Huang, Haizhen Huang, Furu Wei, Feng Sun 0008, Qi Zhang 0066, Zhi Jin 0001 |
EMNLP | 4 |
| 2025 | VirB: A Virus Hierarchical Classification Method Based on ModernBERT
Haizhen Huang, Haodi Feng, Daming Zhu |
ICIC (26) | 1 |
| 2024 | Text Diffusion with Reinforced ConditioningabstractDiffusion models have demonstrated exceptional capability in generating high-quality images, videos, and audio. Due to their adaptiveness in iterative refinement, they provide a strong potential for achieving better non-autoregressive sequence generation. However, existing text diffusion models still fall short in their performance due to a challenge in handling the discreteness of language. This paper thoroughly analyzes text diffusion models and uncovers two significant limitations: degradation of self-conditioning during training and misalignment between training and sampling. Motivated by our findings, we propose a novel Text Diffusion model called TReC, which mitigates the degradation with Reinforced Conditioning and the misalignment by Time-Aware Variance Scaling. Our extensive experiments demonstrate the competitiveness of TReC against autoregressive, non-autoregressive, and diffusion baselines. Moreover, qualitative analysis shows its advanced ability to fully utilize the diffusion process in refining samples. Yuxuan Liu 0011, Tianchi Yang, Shaohan Huang, Haizhen Huang, Furu Wei, Feng Sun 0008, Qi Zhang 0066 |
AAAI | 5 |
| 2024 | HD-Eval: Aligning Large Language Model Evaluators Through Hierarchical Criteria DecompositionabstractYuxuan Liu, Tianchi Yang, Shaohan Huang, Zihan Zhang, Haizhen Huang, Furu Wei, Weiwei Deng, Feng Sun, Qi Zhang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Yuxuan Liu 0011, Tianchi Yang, Shaohan Huang, Haizhen Huang, Furu Wei, Feng Sun 0008, Qi Zhang 0066 |
ACL (1) | 5 |
| 2024 | Calibrating LLM-Based EvaluatorabstractRecent advancements in large language models (LLMs) and their emergent capabilities make LLM a promising reference-free evaluator on the quality of natural language generation, and a competent alternative to human evaluation. However, hindered by the closed-source or high computational demand to host and tune, there is a lack of practice to further calibrate an off-the-shelf LLM-based evaluator towards better human alignment. In this work, we propose AutoCalibrate, a multi-stage, gradient-free approach to automatically calibrate and align an LLM-based evaluator toward human preference. Instead of explicitly modeling human preferences, we first implicitly encompass them within a set of human labels. Then, an initial set of scoring criteria is drafted by the language model itself, leveraging in-context learning on different few-shot examples. To further calibrate this set of criteria, we select the best performers and re-draft them with self-refinement. Our experiments on multiple text quality evaluation datasets illustrate a significant improvement in correlation with expert evaluation through calibration. Our comprehensive qualitative analysis conveys insightful intuitions and observations on the essence of effective scoring criteria. Yuxuan Liu 0011, Tianchi Yang, Shaohan Huang, Haizhen Huang, Furu Wei, Feng Sun 0008, Qi Zhang 0066 |
LREC/COLING | 5 |
| 2023 | Dual-Alignment Pre-training for Cross-lingual Sentence EmbeddingabstractZiheng Li, Shaohan Huang, Zihan Zhang, Zhi-Hong Deng, Qiang Lou, Haizhen Huang, Jian Jiao, Furu Wei, Weiwei Deng, Qi Zhang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Ziheng Li 0003, Shaohan Huang, Zhi-Hong Deng 0001, Qiang Lou, Haizhen Huang, Jian Jiao 0007, Furu Wei, Qi Zhang 0066 |
ACL (1) | 6 |
| 2023 | Towards Better Entity Linking with Multi-View Enhanced DistillationabstractYi Liu, Yuan Tian, Jianxun Lian, Xinlong Wang, Yanan Cao, Fang Fang, Wen Zhang, Haizhen Huang, Weiwei Deng, Qi Zhang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Yi Liu 0067, Jianxun Lian, Yanan Cao 0001, Fang Fang 0009, Haizhen Huang, Qi Zhang 0066 |
ACL (1) | 8 |
| 2023 | Democratizing Reasoning Ability: Tailored Learning from Large Language ModelabstractZhaoyang Wang, Shaohan Huang, Yuxuan Liu, Jiahai Wang, Minghui Song, Zihan Zhang, Haizhen Huang, Furu Wei, Weiwei Deng, Feng Sun, Qi Zhang. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Shaohan Huang, Yuxuan Liu 0011, Jiahai Wang, Minghui Song, Haizhen Huang, Furu Wei, Feng Sun 0008, Qi Zhang 0066 |
EMNLP | 7 |
| 2023 | PASS: Personalized Advertiser-aware Sponsored SearchabstractThe nucleus of online sponsored search systems lies in measuring the relevance between the search intents of users and the advertising purposes of advertisers. Existing conventional doublet-based (query-keyword) relevance models solely rely on short queries and keywords to uncover such intents, which ignore the diverse and personalized preferences of participants (i.e., users and advertisers), resulting in undesirable advertising performance. In this paper, we investigate the novel problem of Personalized A dvertiser-aware Sponsored Search (PASS). Our motivation lies in incorporating the portraits of users and advertisers into relevance models to facilitate the modeling of intrinsic search intents and advertising purposes, leading to a quadruple-based (i.e., user-query-keyword-advertiser) task. Various types of historical behaviors are explored in the format of hypergraphs to provide abundant signals on identifying the preferences of participants. A novel heterogeneous textual hypergraph transformer is further proposed to deeply fuse the textual semantics and the high-order hypergraph topology. Our proposal is extensively evaluated over real industry datasets, and experimental results demonstrate its superiority. Zhoujin Tian, Chaozhuo Li, Zhiqiang Zuo 0004, Zengxuan Wen, Lichao Sun 0001, Xinyue Hu 0003, Haizhen Huang, Senzhang Wang, Xing Xie 0001, Qi Zhang 0066 |
KDD | 8 |
| 2023 | Multi-Grained Topological Pre-Training of Language Models in Sponsored SearchabstractRelevance models measure the semantic closeness between queries and the candidate ads, widely recognized as the nucleus of sponsored search systems. Conventional relevance models solely rely on the textual data within the queries and ads, whose performance is hindered by the scarce semantic information in these short texts. Recently, user behavior graphs have been incorporated to provide complementary information beyond pure textual semantics.Despite the promising performance, behavior-enhanced models suffer from exhausting resource costs due to the extra computations introduced by explicit topological aggregations. In this paper, we propose a novel Multi-Grained Topological Pre-Training paradigm, MGTLM, to teach language models to understand multi-grained topological information in behavior graphs, which contributes to eliminating explicit graph aggregations and avoiding information loss. Extensive experimental results over online and offline settings demonstrate the superiority of our proposal. Zhoujin Tian, Chaozhuo Li, Zhiqiang Zuo 0004, Zengxuan Wen, Xinyue Hu 0003, Haizhen Huang, Senzhang Wang, Xing Xie 0001, Qi Zhang 0066 |
SIGIR | 7 |
| 2022 | PromptBERT: Improving BERT Sentence Embeddings with PromptsabstractTing Jiang, Jian Jiao, Shaohan Huang, Zihan Zhang, Deqing Wang, Fuzhen Zhuang, Furu Wei, Haizhen Huang, Denvy Deng, Qi Zhang. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Jian Jiao 0007, Shaohan Huang, Fuzhen Zhuang, Furu Wei, Haizhen Huang, Denvy Deng, Qi Zhang 0066 |
EMNLP | 8 |
| 2022 | RAPO: An Adaptive Ranking Paradigm for Bilingual Lexicon InductionabstractZhoujin Tian, Chaozhuo Li, Shuo Ren, Zhiqiang Zuo, Zengxuan Wen, Xinyue Hu, Xiao Han, Haizhen Huang, Denvy Deng, Qi Zhang, Xing Xie. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Zhoujin Tian, Chaozhuo Li, Shuo Ren 0002, Zhiqiang Zuo 0004, Zengxuan Wen, Xinyue Hu 0003, Haizhen Huang, Denvy Deng, Qi Zhang 0066, Xing Xie 0001 |
EMNLP | 8 |