Hailong Cao

dblp:43/9013 · DBLP profile ↗
← Back
24ranked-venue papers
6as first author
12since 2021 · last 2026
0000-0002-6842-8674ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 23 · 6 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 4 since 2021
YearPublicationVenuePosition
2026 Long-form RewardBench: Evaluating Reward Models for Long-form Generation
abstract
The widespread adoption of reinforcement learning-based alignment highlights the growing importance of reward models. Various benchmarks have been built to evaluate reward models in various domains and scenarios. However, a significant gap remains in assessing reward models for long-form generation, despite its critical role in real-world applications. To bridge this, we introduce Long-form RewardBench, the first reward modeling testbed specifically designed for long-form generation. Our benchmark encompasses five key subtasks: QA, RAG, Chat, Writing, and Reasoning. We collected instruction and preference data through a meticulously designed multi-stage data collection process, and conducted extensive experiments on 20+ mainstream reward models, including both classifiers and generative models. Our findings reveal that current models still lack long-form reward modeling capabilities. Furthermore, we designed a novel Long-form Needle-in-a-Haystack Test, which revealed a correlation between reward modeling performance and the error's position within a response, as well as the overall response length, with distinct characteristics observed between classification and generative models. Finally, we demonstrate that classifier exhibit better generalizability compared to generative models trained on the same data. As the first benchmark for long-form reward modeling, this work aims to offer a robust platform for visualizing progress in this crucial area.
Hui Huang 0021, Yancheng He, Muyun Yang, Kehai Chen, Conghui Zhu, Hailong Cao, Tiejun Zhao
AAAI9
2026 Lost in Benchmarks? Rethinking Large Language Model Benchmarking with Item Response Theory
abstract
The evaluation of large language models (LLMs) via benchmarks is widespread, yet inconsistencies between different leaderboards and poor separability among top models raise concerns about their ability to accurately reflect authentic model capabilities. This paper provides a critical analysis of benchmark effectiveness, examining mainstream prominent LLM benchmarks using results from diverse models. We first propose Pseudo-Siamese Network for Item Response Theory (PSN-IRT), an enhanced Item Response Theory framework that incorporates a rich set of item parameters within an IRT-grounded architecture. PSN-IRT can be utilized for accurate and reliable estimations of item characteristics and model abilities. Based on PSN-IRT, we conduct extensive analysis on 11 LLM benchmarks comprising 41,871 items, revealing significant and varied shortcomings in their measurement quality. Furthermore, we demonstrate that leveraging PSN-IRT is able to construct smaller benchmarks while maintaining stronger alignment with human preference.
Hongli Zhou 0001, Hui Huang 0021, Ziqing Zhao, Lvyuan Han, Huicheng Wang, Kehai Chen, Muyun Yang, Conghui Zhu, Hailong Cao, Tiejun Zhao
AAAI12
2025 Word-level Cross-lingual Structure in Large Language Models
abstract
Large Language Models (LLMs) have demonstrated exceptional performance across a broad spectrum of cross-lingual Natural Language Processing (NLP) tasks. However, previous methods predominantly focus on leveraging parallel corpus to conduct instruction data for continuing pre-training or fine-tuning. They ignored the state of parallel data on the hidden layers of LLMs. In this paper, we demonstrate Word-level Cross-lingual Structure (WCS) of LLM which proves that the word-level embedding on the hidden layers are isomorphic between languages. We find that the hidden states of different languages’ input on the LLMs hidden layers can be aligned with an orthogonal matrix on word-level. We prove this conclusion in both mathematical and downstream task ways on two representative LLM foundations, LLaMA2 and BLOOM. Besides, we propose an Isomorphism-based Data Augmentation (IDA) method to apply the WCS on a downstream cross-lingual task, Bilingual Lexicon Induction (BLI), in both supervised and unsupervised ways. The experiment shows the significant improvement of our proposed method over all the baselines, especially on low-resource languages.
Hailong Cao, Tiejun Zhao
COLING2
2025 A Chain-of-Task Framework for Instruction Tuning of LLMs Based on Chinese Grammatical Error Correction
abstract
Over-correction is a critical issue for large language models (LLMs) to address Grammatical Error Correction (GEC) task, esp. for Chinese. This paper proposes a Chain-of-Task (CoTask) framework to reduce over-correction. The CoTask framework is applied as multi-task instruction tuning of LLMs by decomposing the process of grammatical error analysis to design auxiliary tasks and adjusting the types and combinations of training tasks. A supervised fine-tuning (SFT) strategy is also presented to enhance the performance of LLMs, together with an algorithm for automatic dataset annotation to avoid additional manual costs. Experimental results demonstrate that our method achieves new state-of-the-art results on both FCGEC (in-domain) and NaCGEC (out-of-domain) test sets.
Xinpeng Liu 0008, Muyun Yang, Hailong Cao, Conghui Zhu, Tiejun Zhao, Wenpeng Lu
COLING4
2025 Self-Relevance-Based Multimodal In-Context Learning for Multimodal Named Entity Recognition
abstract
Recently, Multimodal Named Entity Recognition (MNER) has attracted significant attention. Although MNER utilizing in-context learning has shown improved performance, modality retrieval bias often diminishes the relevance of in-context examples. To address this issue, we propose a self-relevance-based multimodal in-context learning method to mitigate modality retrieval bias by dynamically adjusting the weight of each modality. Specifically, we first measure the self-relevance of the query by calculating the similarity between textual and visual modalities, which helps to assess how much visual information contributes to the textual context. Then, we rank the similarity of different modalities, adjust the image rankings based on self-relevance to reduce modality retrieval bias, and integrate them to select the k most relevant examples. Finally, we use task definition and retrieved examples as effective guidance provided to the Multimodal Large Language Models to obtain feedback. Experimental results demonstrate that our method achieves SOTA performance on two benchmark datasets.
Muyun Yang, Hailong Cao, Conghui Zhu, Wenpeng Lu, Tiejun Zhao
ICME4
2025 Thoughts Behind Attack: Enhancing Security Against Jailbreak Attacks Using Chain-of-Thought
Zhe Tao, Muyun Yang, Hongjiao Guan, Wenpeng Lu, Hailong Cao, Conghui Zhu, Tiejun Zhao
NLPCC (4)6
2025 Enhancing bilingual lexicon induction via harnessing polysemous words
Qiuyu Ding, Hailong Cao, Muyun Yang, Tiejun Zhao
Neurocomputing2
2025 Enhancing word distinction for bilingual lexicon induction with generalized antonym knowledge
Qiuyu Ding, Hailong Cao, Tiejun Zhao
Knowl. Based Syst.2
2024 Enhancing Bilingual Lexicon Induction via Bi-directional Translation Pair Retrieving
abstract
Most Bilingual Lexicon Induction (BLI) methods retrieve word translation pairs by finding the closest target word for a given source word based on cross-lingual word embeddings (WEs). However, we find that solely retrieving translation from the source-to-target perspective leads to some false positive translation pairs, which significantly harm the precision of BLI. To address this problem, we propose a novel and effective method to improve translation pair retrieval in cross-lingual WEs. Specifically, we consider both source-side and target-side perspectives throughout the retrieval process to alleviate false positive word pairings that emanate from a single perspective. On a benchmark dataset of BLI, our proposed method achieves competitive performance compared to existing state-of-the-art (SOTA) methods. It demonstrates effectiveness and robustness across six experimental languages, including similar language pairs and distant language pairs, under both supervised and unsupervised settings.
Qiuyu Ding, Hailong Cao, Tiejun Zhao
AAAI2
2024 Enhancing isomorphism between word embedding spaces for distant languages bilingual lexicon induction
Qiuyu Ding, Hailong Cao, Tiejun Zhao
Neural Comput. Appl.2
2023 Dual Word Embedding for Robust Unsupervised Bilingual Lexicon Induction
abstract
The word embedding models such as Word2vec and FastText simultaneously learn dual representations of input vectors and output vectors. In contrast, almost all existing unsupervised bilingual lexicon induction (UBLI) methods use only input vectors without utilizing output vectors. In this paper, we propose a novel approach to making full use of both input and output vectors for more robust and strong UBLI. We discover the Common Difference Property that one orthogonal transformation can connect not only the input vectors of two languages but also the output vectors. Therefore, we can learn just one transformation to induce two different dictionaries from the input and output vectors, respectively. Between these two quite different dictionaries, a more accurate lexicon with less noise can be induced by taking the intersection of them in UBLI procedure. Extensive experiments show that our method achieves much more robust and strong results than state-of-the-art methods in distant language pairs, while reserving comparable performances in similar language pairs.
Hailong Cao, Liguo Li, Conghui Zhu, Muyun Yang, Tiejun Zhao
IEEE ACM Trans. Audio Speech Lang. Process.1
2022 Cross-lingual Feature Extraction from Monolingual Corpora for Low-resource Unsupervised Bilingual Lexicon Induction
abstract
Despite their progress in high-resource language settings, unsupervised bilingual lexicon induction (UBLI) models often fail on corpora with low-resource distant language pairs due to insufficient initialization. In this work, we propose a cross-lingual feature extraction (CFE) method to learn the cross-lingual features from monolingual corpora for low-resource UBLI, enabling representations of words with the same meaning leveraged by the initialization step. By integrating cross-lingual representations with pre-trained word embeddings in a fully unsupervised initialization on UBLI, the proposed method outperforms existing state-of-the-art methods on low-resource language pairs (EN-VI, EN-TH, EN-ZH, EN-JA). The ablation study also proves that the learned cross-lingual features can enhance the representational ability and robustness of the existing embedding model.
Hailong Cao, Tiejun Zhao, Wei Peng 0011
COLING2
2019 A Bilingual Adversarial Autoencoder for Unsupervised Bilingual Lexicon Induction
abstract
Unsupervised bilingual lexicon induction aims to generate bilingual lexicons without any cross-lingual signals. Successfully solving this problem would benefit many downstream tasks, such as unsupervised machine translation and transfer learning. In this work, we propose an unsupervised framework, named bilingual adversarial autoencoder, which automatically generates bilingual lexicon for a pair of languages from their monolingual word embeddings. In contrast to existing frameworks which learn a direct cross-lingual mapping of word embeddings from the source language to the target language, we train two autoencoders jointly to transform the source and the target monolingual word embeddings into a shared embedding space, where a word and its translation are close to each other. In this way, we capture the cross-lingual features of word embeddings from different languages and use them to induce bilingual lexicons. By conducting extensive experiments across eight language pairs, we demonstrate that the proposed method significantly outperforms the existing adversarial methods and even achieves best-published results across most language pairs.
Xuefeng Bai 0001, Hailong Cao, Kehai Chen, Tiejun Zhao
IEEE ACM Trans. Audio Speech Lang. Process.2
2018 Point Set Registration for Unsupervised Bilingual Lexicon Induction
abstract
Inspired by the observation that word embeddings exhibit isomorphic structure across languages, we propose a novel method to induce a bilingual lexicon from only two sets of word embeddings, which are trained on monolingual source and target data respectively. This is achieved by formulating the task as point set registration which is a more general problem. We show that a transformation from the source to the target embedding space can be learned automatically without any form of cross-lingual supervision. By properly adapting a traditional point set registration model to make it be suitable for processing word embeddings, we achieved state-of-the-art performance on the unsupervised bilingual lexicon induction task. The point set registration problem has been well-studied and can be solved by many elegant models, we thus opened up a new opportunity to capture the universal lexical semantic structure across languages.
Hailong Cao, Tiejun Zhao
IJCAI1
2018 Improving Vector Space Word Representations Via Kernel Canonical Correlation Analysis
abstract
Cross-lingual word embeddings are representations for vocabularies of two or more languages in one common continuous vector space and are widely used in various natural language processing tasks. A state-of-the-art way to generate cross-lingual word embeddings is to learn a linear mapping, with an assumption that the vector representations of similar words in different languages are related by a linear relationship. However, this assumption does not always hold true, especially for substantially different languages. We therefore propose to use kernel canonical correlation analysis to capture a non-linear relationship between word embeddings of two languages. By extensively evaluating the learned word embeddings on three tasks (word similarity, cross-lingual dictionary induction, and cross-lingual document classification) across five language pairs, we demonstrate that our proposed approach achieves essentially better performances than previous linear methods on all of the three tasks, especially for language pairs with substantial typological difference.
Xuefeng Bai 0001, Hailong Cao, Tiejun Zhao
ACM Trans. Asian Low Resour. Lang. Inf. Process.2
2016 A Distribution-based Model to Learn Bilingual Word Embeddings
abstract
We introduce a distribution based model to learn bilingual word embeddings from monolingual data. It is simple, effective and does not require any parallel data or any seed lexicon. We take advantage of the fact that word embeddings are usually in form of dense real-valued low-dimensional vector and therefore the distribution of them can be accurately estimated. A novel cross-lingual learning objective is proposed which directly matches the distributions of word embeddings in one language with that in the other language. During the joint learning process, we dynamically estimate the distributions of word embeddings in two languages respectively and minimize the dissimilarity between them through standard back propagation algorithm. Our learned bilingual word embeddings allow to group each word and its translations together in the shared vector space. We demonstrate the utility of the learned embeddings on the task of finding word-to-word translations from monolingual corpora. Our model achieved encouraging performance on data in both related languages and substantially different languages.
Hailong Cao, Tiejun Zhao
COLING1
2016 Improving Dependency Parsing on Clinical Text with Syntactic Clusters from Web Text
Xiuming Qiao, Hailong Cao, Tiejun Zhao, Kehai Chen
ICONIP (1)2
2016 Improving Unsupervised Dependency Parsing with Knowledge from Query Logs
abstract
Unsupervised dependency parsing becomes more and more popular in recent years because it does not need expensive annotations, such as treebanks, which are required for supervised and semi-supervised dependency parsing. However, its accuracy is still far below that of supervised dependency parsers, partly due to the fact that their parsing model is insufficient to capture linguistic phenomena underlying texts. The performance for unsupervised dependency parsing can be improved by mining knowledge from the texts and by incorporating it into the model. In this article, syntactic knowledge is acquired from query logs to help estimate better probabilities in dependency models with valence. The proposed method is language independent and obtains an improvement of 4.1% unlabeled accuracy on the Penn Chinese Treebank by utilizing additional dependency relations from the Sogou query logs and Baidu query logs. Morever, experiments show that the proposed model achieves improvements of 8.07% on CoNLL 2007 English using the AOL query logs. We believe query logs are useful sources of syntactic knowledge for many natural language processing (NLP) tasks.
Xiuming Qiao, Hailong Cao, Tiejun Zhao
ACM Trans. Asian Low Resour. Lang. Inf. Process.2
2014 A Lexicalized Reordering Model for Hierarchical Phrase-based Translation
Hailong Cao, Dongdong Zhang 0001, Mu Li 0001, Ming Zhou 0001, Tiejun Zhao
COLING1
2014 Soft Dependency Matching for Hierarchical Phrase-based Machine Translation
Hailong Cao, Dongdong Zhang 0001, Ming Zhou 0001, Tiejun Zhao
COLING1
2014 Discriminative Training for Log-Linear Based SMT: Global or Local Methods
abstract
In statistical machine translation, the standard methods such as MERT tune a single weight with regard to a given development data. However, these methods suffer from two problems due to the diversity and uneven distribution of source sentences. First, their performance is highly dependent on the choice of a development set, which may lead to an unstable performance for testing. Second, the sentence level translation quality is not assured since tuning is performed on the document level rather than on sentence level. In contrast with the standard global training in which a single weight is learned, we propose novel local training methods to address these two problems. We perform training and testing in one step by locally learning the sentence-wise weight for each input sentence. Since the time of each tuning step is unnegligible and learning sentence-wise weights for the entire test set means many passes of tuning, it is a great challenge for the efficiency of local training. We propose an efficient two-phase method to put the local training into practice by employing the ultraconservative update. On NIST Chinese-to-English translation tasks with both medium and large scales of training data, our local training methods significantly outperform standard methods with the maximal improvements up to 2.0 BLEU points, meanwhile their efficiency is comparable to that of the standard methods.
Lemao Liu, Tiejun Zhao, Taro Watanabe, Hailong Cao, Conghui Zhu
ACM Trans. Asian Lang. Inf. Process.4
2012 Phrasal Syntactic Category Sequence Model for Phrase-Based MT
Hailong Cao, Eiichiro Sumita, Tiejun Zhao, Sheng Li 0003
CICLing (2)1
2012 Locally Training the Log-Linear Model for SMT
Lemao Liu, Hailong Cao, Taro Watanabe, Tiejun Zhao, Mo Yu, Conghui Zhu
EMNLP-CoNLL2
2011 A Unified and Discriminative Soft Syntactic Constraint Model for Hierarchical Phrase-based Translation
Lemao Liu, Tiejun Zhao, Hailong Cao
MTSummit4