VLDB 2026 Research / reviewers in the wild / expert
Hui Huang 0021
dblp:33/5763-21
· DBLP profile ↗
11ranked-venue papers
5as first author
11since 2021 · last 2026
0000-0002-3544-9719ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 5 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Long-form RewardBench: Evaluating Reward Models for Long-form GenerationabstractThe widespread adoption of reinforcement learning-based alignment highlights the growing importance of reward models. Various benchmarks have been built to evaluate reward models in various domains and scenarios. However, a significant gap remains in assessing reward models for long-form generation, despite its critical role in real-world applications. To bridge this, we introduce Long-form RewardBench, the first reward modeling testbed specifically designed for long-form generation. Our benchmark encompasses five key subtasks: QA, RAG, Chat, Writing, and Reasoning. We collected instruction and preference data through a meticulously designed multi-stage data collection process, and conducted extensive experiments on 20+ mainstream reward models, including both classifiers and generative models. Our findings reveal that current models still lack long-form reward modeling capabilities. Furthermore, we designed a novel Long-form Needle-in-a-Haystack Test, which revealed a correlation between reward modeling performance and the error's position within a response, as well as the overall response length, with distinct characteristics observed between classification and generative models. Finally, we demonstrate that classifier exhibit better generalizability compared to generative models trained on the same data. As the first benchmark for long-form reward modeling, this work aims to offer a robust platform for visualizing progress in this crucial area. Hui Huang 0021, Yancheng He, Muyun Yang, Kehai Chen, Conghui Zhu, Hailong Cao, Tiejun Zhao |
AAAI | 1 |
| 2026 | Think-J: Learning to Think for Generative LLM-as-a-JudgeabstractLLM-as-a-Judge refers to the automatic modeling of preferences for responses generated by Large Language Models (LLMs), which is of significant importance for both LLM evaluation and reward modeling. Although generative LLMs have made substantial progress in various tasks, their performance as LLM-Judge still falls short of expectations. In this work, we propose Think-J, which improves generative LLM-as-a-Judge by learning how to think. We first utilized a small amount of curated data to develop the model with initial judgment thinking capabilities. Subsequently, we optimize the judgment thinking traces based on reinforcement learning (RL). We propose two methods for judgment thinking optimization, based on offline and online RL, respectively. The offline method requires training a critic model to construct positive and negative examples for learning. The online method defines rule-based reward as feedback for optimization. Experimental results showed that our approach can significantly enhance the evaluation capability of generative LLM-Judge, surpassing both generative and classifier-based LLM-Judge without requiring extra human annotations. Hui Huang 0021, Yancheng He, Hongli Zhou 0001, Weixun Wang, Wenbo Su |
AAAI | 1 |
| 2026 | Lost in Benchmarks? Rethinking Large Language Model Benchmarking with Item Response TheoryabstractThe evaluation of large language models (LLMs) via benchmarks is widespread, yet inconsistencies between different leaderboards and poor separability among top models raise concerns about their ability to accurately reflect authentic model capabilities. This paper provides a critical analysis of benchmark effectiveness, examining mainstream prominent LLM benchmarks using results from diverse models. We first propose Pseudo-Siamese Network for Item Response Theory (PSN-IRT), an enhanced Item Response Theory framework that incorporates a rich set of item parameters within an IRT-grounded architecture. PSN-IRT can be utilized for accurate and reliable estimations of item characteristics and model abilities. Based on PSN-IRT, we conduct extensive analysis on 11 LLM benchmarks comprising 41,871 items, revealing significant and varied shortcomings in their measurement quality. Furthermore, we demonstrate that leveraging PSN-IRT is able to construct smaller benchmarks while maintaining stronger alignment with human preference. Hongli Zhou 0001, Hui Huang 0021, Ziqing Zhao, Lvyuan Han, Huicheng Wang, Kehai Chen, Muyun Yang, Conghui Zhu, Hailong Cao, Tiejun Zhao |
AAAI | 2 |
| 2025 | Chinese SimpleQA: A Chinese Factuality Evaluation for Large Language ModelsabstractYancheng He, Shilong Li, Jiaheng Liu, Yingshui Tan, Weixun Wang, Hui Huang, Xingyuan Bu, Hangyu Guo, Chengwei Hu, Boren Zheng, Zhuoran Lin, Dekai Sun, Zhicheng Zheng, Wenbo Su, Bo Zheng. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Yancheng He, Yingshui Tan, Weixun Wang, Hui Huang 0021, Xingyuan Bu, Hangyu Guo, Chengwei Hu, Boren Zheng, Zhuoran Lin, Dekai Sun, Zhicheng Zheng, Wenbo Su, Bo Zheng 0007 |
ACL (1) | 6 |
| 2025 | MuSC: Improving Complex Instruction Following with Multi-granularity Self-Contrastive TrainingabstractHui Huang, Jiaheng Liu, Yancheng He, Shilong Li, Bing Xu, Conghui Zhu, Muyun Yang, Tiejun Zhao. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Hui Huang 0021, Yancheng He, Conghui Zhu, Muyun Yang, Tiejun Zhao |
ACL (1) | 1 |
| 2025 | SCM: Enhancing Large Language Model with Self-Controlled Memory Framework
Xinnian Liang, Jian Yang 0003, Hui Huang 0021, Zhenhe Wu, Shuangzhi Wu, Zejun Ma 0001, Zhoujun Li 0001 |
DASFAA (6) | 4 |
| 2025 | AIR: Complex Instruction Generation via Automatic Iterative RefinementabstractWith the development of large language models, their ability to follow simple instructions has significantly improved.However, adhering to complex instructions remains a major challenge.Current approaches to generating complex instructions are often irrelevant to the current instruction requirements or suffer from limited scalability and diversity.Moreover, methods such as back-translation, while effective for simple instruction generation, fail to leverage the rich knowledge and formatting in human written documents.In this paper, we propose a novel Automatic Iterative Refinement (AIR) framework to generate complex instructions with constraints, which not only better reflects the requirements of real scenarios but also significantly enhances LLMs' ability to follow complex instructions.The AIR framework consists of two stages: 1) Generate an initial instruction from a document; 2) Iteratively refine instructions with LLM-as-judge guidance by comparing the model's output with the document to incorporate valuable constraints.Finally, we construct the AIR-10K dataset with 10K complex instructions and demonstrate that instructions generated with our approach significantly improve the model's ability to follow complex instructions, outperforming existing methods for instruction generation 1 . Model Answer Refined D Model Answer Check Constraints Identify ConstraintsI: Write a casual review of a water-proof camera. Yancheng He, Yu Li 0007, Hui Huang 0021, Chengwei Hu, Wenbo Su, Bo Zheng 0007 |
EMNLP | 4 |
| 2025 | Legal Fact Prediction: The Missing Piece in Legal Judgment PredictionabstractJunkai Liu, Yujie Tong, Hui Huang, Bowen Zheng, Yiran Hu, Peicheng Wu, Chuan Xiao, Makoto Onizuka, Muyun Yang, Shuyuan Zheng. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Yujie Tong, Hui Huang 0021, Yiran Hu, Peicheng Wu, Chuan Xiao 0001, Makoto Onizuka, Muyun Yang, Shuyuan Zheng |
EMNLP | 3 |
| 2025 | DREAM: Disentangling Risks to Enhance Safety Alignment in Multimodal Large Language ModelsabstractJianyu Liu, Hangyu Guo, Ranjie Duan, Xingyuan Bu, Yancheng He, Shilong Li, Hui Huang, Jiaheng Liu, Yucheng Wang, Chenchen Jing, Xingwei Qu, Xiao Zhang, Pei Wang, Yanan Wu, Jihao Gu, Yangguang Li, Jianke Zhu. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Jianyu Liu, Hangyu Guo, Ranjie Duan, Xingyuan Bu, Yancheng He, Hui Huang 0021, Chenchen Jing, Xingwei Qu, Jihao Gu, Yangguang Li 0001, Jianke Zhu |
NAACL (Long Papers) | 7 |
| 2023 | Improving Translation Quality Estimation with Bias MitigationabstractState-of-the-art translation Quality Estimation (QE) models are proven to be biased.More specifically, they over-rely on monolingual features while ignoring the bilingual semantic alignment.In this work, we propose a novel method to mitigate the bias of the QE model and improve estimation performance.Our method is based on the contrastive learning between clean and noisy sentence pairs.We first introduce noise to the target side of the parallel sentence pair, forming the negative samples.With the original parallel pairs as the positive sample, the QE model is contrastively trained to distinguish the positive samples from the negative ones.This objective is jointly trained with the regression-style quality estimation, so as to prevent the QE model from overfitting to monolingual features.Experiments on WMT QE evaluation datasets demonstrate that our method improves the estimation performance by a large margin while mitigating the bias 1 . Hui Huang 0021, Shuangzhi Wu, Kehai Chen, Hui Di, Muyun Yang, Tiejun Zhao |
ACL (1) | 1 |
| 2023 | Towards Making the Most of LLM for Translation Quality Estimation
Hui Huang 0021, Shuangzhi Wu, Xinnian Liang, Yanrui Shi, Peihao Wu, Muyun Yang, Tiejun Zhao |
NLPCC (1) | 1 |