Jiahuan Li

dblp:199/0384 · DBLP profile ↗
← Back
13ranked-venue papers
6as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 5 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Language models and text generation · 37% Representation and self-supervised learning · 15% Machine translation · 10%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational finance and economics · 100%

Topics — the 18 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Machine translation › machine translation evaluation
translation quality estimation
1.222023
Denoising Pre-training for Machine Translation Quality Estimation with Curriculum Learning · AAAI 2023
DirectQE: Direct Pretraining for Machine Translation Quality Estimation · AAAI 2021
Machine learning › Generative modeling
synthetic data generation
1.012026
Unlocking Implicit Experience: Synthesizing Tool-Use Trajectories from Text · ACL (1) 2026
Natural language and speech › Language models and text generation › agentic language model
tool-augmented language models
1.012026
Unlocking Implicit Experience: Synthesizing Tool-Use Trajectories from Text · ACL (1) 2026
Machine learning › Representation and self-supervised learning
pre-training
0.932024
DirectQE: Direct Pretraining for Machine Translation Quality Estimation · AAAI 2021
PreAlign: Boosting Cross-Lingual Transfer by Early Establishment of Multilingual Alignment · EMNLP 2024
Denoising Pre-training for Machine Translation Quality Estimation with Curriculum Learning · AAAI 2023
Computer vision › Vision and language
multimodal benchmark
0.912025
VisFinEval: A Scenario-Driven Chinese Multimodal Benchmark for Holistic Financial Understanding · EMNLP 2025
Computational finance and economics › financial data analysis
financial document analysis
0.912025
VisFinEval: A Scenario-Driven Chinese Multimodal Benchmark for Holistic Financial Understanding · EMNLP 2025
Machine learning › Representation and self-supervised learning › representation matching › feature alignment › embedding alignment
cross-lingual alignment
0.812024
PreAlign: Boosting Cross-Lingual Transfer by Early Establishment of Multilingual Alignment · EMNLP 2024
Machine learning › Transfer learning and domain adaptation
cross-lingual transfer
0.812024
PreAlign: Boosting Cross-Lingual Transfer by Early Establishment of Multilingual Alignment · EMNLP 2024
Natural language and speech › Language models and text generation › retrieval-augmented generation
knowledge conflict
0.812024
Formality is Favored: Unraveling the Learning Preferences of Large Language Models on Data with Conflicting Knowledge · EMNLP 2024
Natural language and speech › Language models and text generation
multilingual language models
0.812024
PreAlign: Boosting Cross-Lingual Transfer by Early Establishment of Multilingual Alignment · EMNLP 2024
Machine learning › Reinforcement learning
preference learning
0.812024
Formality is Favored: Unraveling the Learning Preferences of Large Language Models on Data with Conflicting Knowledge · EMNLP 2024
Natural language and speech › Language models and text generation › large language model training
pretraining data quality
0.812024
Formality is Favored: Unraveling the Learning Preferences of Large Language Models on Data with Conflicting Knowledge · EMNLP 2024
Natural language and speech › Language models and text generation
pseudo data generation
0.512021
DirectQE: Direct Pretraining for Machine Translation Quality Estimation · AAAI 2021
Natural language and speech › Language models and text generation › text generation › conditional text generation
definition generation
0.412020
Explicit Semantic Decomposition for Definition Generation · ACL 2020
Computer vision › Segmentation and scene understanding
semantic decomposition
0.412020
Explicit Semantic Decomposition for Definition Generation · ACL 2020
Natural language and speech › Question answering and dialogue systems
knowledge-intensive tasks
0.212024
Formality is Favored: Unraveling the Learning Preferences of Large Language Models on Data with Conflicting Knowledge · EMNLP 2024
Machine learning › Learning paradigms › curriculum learning
curriculum learning for pre-training
0.212023
Denoising Pre-training for Machine Translation Quality Estimation with Curriculum Learning · AAAI 2023
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
discrete latent variable
0.112020
Explicit Semantic Decomposition for Definition Generation · ACL 2020

Methods — techniques the papers use, named apart from their topics

multimodal large language model evaluation · 1.7supervised fine-tuning · 1.0large language model · 1.0statistical co-occurrence analysis · 0.8representation initialization · 0.8code-switching · 0.8statistical and distributional noise metrics · 0.7denoising pretraining · 0.7curriculum learning · 0.7predictor-estimator framework · 0.5
YearPublicationVenuePosition
2026 Unlocking Implicit Experience: Synthesizing Tool-Use Trajectories from Text
abstract
Zhihao Xu, Rumei Li, Jiahuan Li, Rongxiang Weng, Jingang Wang, Xunliang Cai, Xiting Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Jiahuan Li, Rongxiang Weng, Jingang Wang, Xiting Wang
ACL (1)3
2026 Why not transform chat large language models to non-English?
Xiang Geng, Ming Zhu 0010, Jiahuan Li, Zhejian Lai, Shuaijie She, Yinglu Li, Yuang Li, Chang Su 0001, Xinglin Lyu, Min Zhang 0042, Jiajun Chen 0001, Hao Yang 0006, Shujian Huang
Frontiers Comput. Sci.3
2025 VisFinEval: A Scenario-Driven Chinese Multimodal Benchmark for Holistic Financial Understanding
abstract
Zhaowei Liu, Xin Guo, Haotian Xia, Lingfeng Zeng, Fangqi Lou, Jinyi Niu, Mengping Li, Qi Qi, Jiahuan Li, Wei Zhang, Yinglong Wang, Weige Cai, Weining Shen, Liwen Zhang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Haotian Xia, Lingfeng Zeng, Fangqi Lou, Jinyi Niu, Mengping Li, Jiahuan Li, Weige Cai, Weining Shen
EMNLP9
2025 "I've Heard of You!": Generate Spoken Named Entity Recognition Data for Unseen Entities
abstract
Spoken named entity recognition (NER) aims to identify named entities from speech, playing an important role in speech processing. New named entities appear every day, however, annotating their Spoken NER data is costly. In this paper, we demonstrate that existing Spoken NER systems perform poorly when dealing with previously unseen named entities. To tackle this challenge, we propose a method for generating Spoken NER data based on a named entity dictionary (NED) to reduce costs. Specifically, we first use a large language model (LLM) to generate sentences from the sampled named entities and then use a text-to-speech (TTS) system to generate the speech. Furthermore, we introduce a noise metric to filter out noisy data. To evaluate our approach, we release a novel Spoken NER benchmark along with a corresponding NED containing 8,853 entities. Experiment results show that our method achieves state-of-the-art (SOTA) performance in the in-domain, zero-shot domain adaptation, and fully zero-shot settings. Our data will be available at https://github.com/DeepLearnXMU/HeardU.
Xiang Geng, Yuang Li, Mengxin Ren, Wei Tang 0013, Jiahuan Li, Zhibin Lan, Min Zhang 0042, Hao Yang 0006, Shujian Huang, Jinsong Su
ICASSP6
2025 Wavelength- and Depth-Aware Deep Image Prior for Blind Hyperspectral Imagery Deblurring with Coarse Depth Guidance
abstract
Hyperspectral imagery (HSI) provides detailed spectral information, enabling precise analysis of materials. However, HSI imaging suffers from blurring degradation which results in the loss of fine details and hinders subsequent applications. The degree of blurriness is highly related to wavelength and depth, existing deblurring methods either lack the utilization of spectral correlation or ignore the depth variation since paired HSI and depth data are difficult to acquire and less discussed, leading to degraded performance when encountering wide-range HSIs of non-planar scenes. To address these challenges in both data acquisition and algorithm design, we propose a novel approach that simultaneously collects both modalities and integrates depth refinement into a blind HSI deblurring model with wavelength- and depth-aware deep image prior. Specifically, we capture blurred HSI and coarse depth map with separate devices, followed by registration. Our method performs depth-guided deblurring through depth-variant multi-channel kernel estimation and soft-weight map-based layer composition, while simultaneously refining the depth. The proposed approach effectively restores fine details with fewer artifacts, showing superior performance for both simulated blurred HSIs and real captured HSIs.
Jiahuan Li, Wei He 0003, Naoto Yokoya
WACV1
2025 Visual primitives as words: Alignment and interaction for compositional zero-shot learning
Feng Shuang 0002, Jiahuan Li, Qingbao Huang, Wenye Zhao, Dongsheng Xu 0001, Haonan Cheng
Pattern Recognit.2
2024 Formality is Favored: Unraveling the Learning Preferences of Large Language Models on Data with Conflicting Knowledge
abstract
Having been trained on massive pretraining data, large language models have shown excellent performance on many knowledge-intensive tasks.However, pretraining data tends to contain misleading and even conflicting information, and it is intriguing to understand how LLMs handle these noisy data during training.In this study, we systematically analyze LLMs' learning preferences for data with conflicting knowledge.We find that pretrained LLMs establish learning preferences similar to humans, i.e., preferences towards formal texts and texts with fewer spelling errors, resulting in faster learning and more favorable treatment of knowledge in data with such features when facing conflicts.This finding is generalizable across models and languages and is more evident in larger models.An in-depth analysis reveals that LLMs tend to trust data with features that signify consistency with the majority of data, and it is possible to instill new preferences and erase old ones by manipulating the degree of consistency with the majority data.
Jiahuan Li, Yiqing Cao, Shujian Huang, Jiajun Chen 0001
EMNLP1
2024 PreAlign: Boosting Cross-Lingual Transfer by Early Establishment of Multilingual Alignment
abstract
Large language models demonstrate reasonable multilingual abilities, despite predominantly English-centric pretraining.However, the spontaneous multilingual alignment in these models is shown to be weak, leading to unsatisfactory cross-lingual transfer and knowledge sharing.Previous works attempt to address this issue by explicitly injecting multilingual alignment information during or after pretraining.Thus for the early stage in pretraining, the alignment is weak for sharing information or knowledge across languages.In this paper, we propose PREALIGN, a framework that establishes multilingual alignment prior to language model pretraining.PREALIGN injects multilingual alignment by initializing the model to generate similar representations of aligned words and preserves this alignment using a code-switching strategy during pretraining.Extensive experiments in a synthetic English to English-Clone setting demonstrate that PREALIGN significantly outperforms standard multilingual joint training in language modeling, zero-shot crosslingual transfer, and cross-lingual knowledge application.Further experiments in real-world scenarios further validate PREALIGN's effectiveness across various languages and model sizes.
Jiahuan Li, Shujian Huang, Aarron Ching, Xinyu Dai, Jiajun Chen 0001
EMNLP1
2024 MT-PATCHER: Selective and Extendable Knowledge Distillation from Large Language Models for Machine Translation
abstract
Jiahuan Li, Shanbo Cheng, Shujian Huang, Jiajun Chen. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Jiahuan Li, Shanbo Cheng, Shujian Huang, Jiajun Chen 0001
NAACL-HLT1
2024 Eliciting the Translation Ability of Large Language Models via Multilingual Finetuning with Translation Instructions
abstract
Abstract Large-scale pretrained language models (LLMs), such as ChatGPT and GPT4, have shown strong abilities in multilingual translation, without being explicitly trained on parallel corpora. It is intriguing how the LLMs obtain their ability to carry out translation instructions for different languages. In this paper, we present a detailed analysis by finetuning a multilingual pretrained language model, XGLM-7.5B, to perform multilingual translation following given instructions. Firstly, we show that multilingual LLMs have stronger translation abilities than previously demonstrated. For a certain language, the translation performance depends on its similarity to English and the amount of data used in the pretraining phase. Secondly, we find that LLMs’ ability to carry out translation instructions relies on the understanding of translation instructions and the alignment among different languages. With multilingual finetuning with translation instructions, LLMs could learn to perform the translation task well even for those language pairs unseen during the instruction tuning phase.
Jiahuan Li, Hao Zhou 0012, Shujian Huang, Shanbo Cheng, Jiajun Chen 0001
Trans. Assoc. Comput. Linguistics1
2023 Denoising Pre-training for Machine Translation Quality Estimation with Curriculum Learning
abstract
Quality estimation (QE) aims to assess the quality of machine translations when reference translations are unavailable. QE plays a crucial role in many real-world applications of machine translation. Because labeled QE data are usually limited in scale, recent research, such as DirectQE, pre-trains QE models with pseudo QE data and obtains remarkable performance. However, there tends to be inevitable noise in the pseudo data, hindering models from learning QE accurately. Our study shows that the noise mainly comes from the differences between pseudo and real translation outputs. To handle this problem, we propose CLQE, a denoising pre-training framework for QE based on curriculum learning. More specifically, we propose to measure the degree of noise in the pseudo QE data with some metrics based on statistical or distributional features. With the guidance of these metrics, CLQE gradually pre-trains the QE model using data from cleaner to noisier. Experiments on various benchmarks reveal that CLQE outperforms DirectQE and other strong baselines. We also show that with our framework, pre-training converges faster than directly using the pseudo data. We make our CLQE code available (https://github.com/NJUNLP/njuqe).
Xiang Geng, Jiahuan Li, Shujian Huang, Hao Yang 0006, Shimin Tao, Jiajun Chen 0001
AAAI3
2021 DirectQE: Direct Pretraining for Machine Translation Quality Estimation
abstract
Machine Translation Quality Estimation (QE) is a task of predicting the quality of machine translations without relying on any reference. Recently, the predictor-estimator framework trains the predictor as a feature extractor, which leverages the extra parallel corpora without QE labels, achieving promising QE performance. However, we argue that there are gaps between the predictor and the estimator in both data quality and training objectives, which preclude QE models from benefiting from a large number of parallel corpora more directly. We propose a novel framework called DirectQE that provides a direct pretraining for QE tasks. In DirectQE, a generator is trained to produce pseudo data that is closer to the real QE data, and a detector is pretrained on these data with novel objectives that are akin to the QE task. Experiments on widely used benchmarks show that DirectQE outperforms existing methods, without using any pretraining models such as BERT. We also give extensive analyses showing how fixing the two gaps contributes to our improvements.
Qu Cui, Shujian Huang, Jiahuan Li, Xiang Geng, Zaixiang Zheng, Guoping Huang, Jiajun Chen 0001
AAAI3
2020 Explicit Semantic Decomposition for Definition Generation
abstract
Definition generation, which aims to automatically generate dictionary definitions for words, has recently been proposed to assist the construction of dictionaries and help people understand unfamiliar texts.However, previous works hardly consider explicitly modeling the "components" of definitions, leading to under-specific generation results.In this paper, we propose ESD, namely Explicit Semantic Decomposition for definition generation, which explicitly decomposes meaning of words into semantic components, and models them with discrete latent variables for definition generation.Experimental results show that ESD achieves substantial improvements on WordNet and Oxford benchmarks over strong previous baselines.
Jiahuan Li, Shujian Huang, Xinyu Dai, Jiajun Chen 0001
ACL1