VLDB 2026 Research / reviewers in the wild / expert
Yichong Huang
dblp:291/4211 · also Yi-Chong Huang
· DBLP profile ↗
13ranked-venue papers
5as first author
13since 2021 · last 2026
0009-0005-4004-8564ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 5 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | The Visual Prism: Refracting Images into Parallel Multilingual Descriptions with Structured Visual GuidanceabstractParallel corpora, as the foundation of machine translation, remain crucial even in the era of large language models (LLMs) for pre-training and fine-tuning. However, annotating parallel corpora is extremely costly, as it requires annotators to be proficient in multiple languages. To reduce this cost, prior work has explored image-pivoted corpus synthesis, generating multilingual captions for the same image as pseudo-parallel data. Unfortunately, these pseudo corpora suffer from the serious issue of multilingual focus divergence, i.e., the model attending to distinct aspects of the image when generating captions in different languages. To address this problem, we propose a method called PRISMS (Parallel Refracting ImageS into Multilingual descriptions with Structured visual guidance), which leverages semantic graphs as structured visual guidance to unify the focus of multilingual captions. To ensure adherence to this guidance, we introduce two key techniques: supervised fine-tuning using self-generated instructional data, and reinforcement learning with a reward signal based on semantic graph consistency. Experimental results on five languages show that our PRISMS significantly improves the image-pivot parallel corpora synthesis, enabling LLMs to achieve translation performance comparable to that of models trained on manually annotated corpora. Chengpeng Fu, Yichong Huang, Wenshuai Huo, Baohang Li, Yang Xiang 0003, Ting Liu 0001 |
AAAI | 3 |
| 2026 | English is not all you need: Rewarding better translation to inspire multilingual capability in LLMs
Wenshuai Huo, Yichong Huang, Chengpeng Fu, Hui Wang 0030, Bing Qin 0001 |
Neurocomputing | 3 |
| 2026 | S HARING B EYOND D ECISION : Deep Collaboration between Large Language Models via Representation EnsembleabstractAbstract Large Language Models (LLMs) exhibit unique strengths arising from differences in model architecture, training data, and strategies. Ensemble learning has been explored to leverage these complementary strengths through decision-level sharing (i.e.,Decision Ensemble), which combines the predictions from multiple LLMs. However, such methods integrate only shallow decisions and overlook the exchange of deeper levels of information within the internal representations of LLMs, such as problem understanding, world knowledge, and latent reasoning patterns. In this work, we propose Representation Ensemble (RISE), a novel ensemble framework that enables cross-LLM representation sharing for richer information exchange. To address challenges of representation-level interaction caused by layer misalignment and latent-space incompatibility across LLMs, we introduce a representation alignment method based on relational similarity measures and an orthogonal latent-space transformation. Experimental results show that (1) RISE achieves performance competitive with existing decision ensemble methods, and (2) RISE is strongly complementary to decision ensemble, with their combination boosting collaboration gains by 14%–41%. Finally, we further compare ensemble of small LLMs to a single larger LLM and to model merging and composition approaches, and find that ensemble learning consistently generalizes well without additional training. Yichong Huang, Jinlan Fu, Xiachong Feng, Baohang Li, Zekai Ye, Libo Qin 0001, Hao Fei 0001, See-Kiong Ng, Bing Qin 0001 |
Trans. Assoc. Comput. Linguistics | 1 |
| 2025 | Enhancing Non-English Capabilities of English-Centric Large Language Models Through Deep Supervision Fine-TuningabstractLarge language models (LLMs) have demonstrated significant progress in multilingual language understanding and generation. However, due to the imbalance in training data, their capabilities in non-English languages are limited. Recent studies revealed the English-pivot multilingual mechanism of LLMs, where LLMs implicitly convert non-English queries into English ones at the bottom layers and adopt English for thinking at the middle layers. However, due to the absence of explicit supervision for cross-lingual alignment in the intermediate layers of LLMs, the internal representations during these stages may become inaccurate. In this work, we introduce a deep supervision fine-tuning method (DFT) that incorporates additional supervision in the internal layers of the model to guide its workflow. Specifically, we introduce two training objectives on different layers of LLMs: one at the bottom layers to constrain the conversion of the target language into English, and another at the middle layers to constrain reasoning in English. To effectively achieve the guiding purpose, we designed two types of supervision signals: logits and feature, which represent a stricter constraint and a relatively more relaxed guidance. Our method guides the model to not only consider the final generated result when processing non-English inputs but also ensure the accuracy of internal representations. We conducted extensive experiments on typical English-centric large models, LLaMA-2 and Gemma-2, and the results on multiple multilingual datasets show that our method significantly outperforms traditional fine-tuning methods. Wenshuai Huo, Yichong Huang, Chengpeng Fu, Baohang Li, Yangfan Ye, Zhirui Zhang, Dandan Tu, Duyu Tang, Yunfei Lu, Hui Wang 0030, Bing Qin 0001 |
AAAI | 3 |
| 2025 | One for All: Update Parameterized Knowledge Across Multiple Models with Once EditabstractWeitao Ma, Xiyuan Du, Xiaocheng Feng, Lei Huang, Yichong Huang, Huiyi Zhang, Xiaoliang Yang, Baohang Li, Xiachong Feng, Ting Liu, Bing Qin. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Weitao Ma, Xiyuan Du, Lei Huang 0021, Yichong Huang, Huiyi Zhang, Xiaoliang Yang, Baohang Li, Xiachong Feng, Ting Liu 0001, Bing Qin 0001 |
ACL (1) | 5 |
| 2025 | CC-Tuning: A Cross-Lingual Connection Mechanism for Improving Joint Multilingual Supervised Fine-TuningabstractYangfan Ye, Xiaocheng Feng, Zekun Yuan, Xiachong Feng, Libo Qin, Lei Huang, Weitao Ma, Yichong Huang, Zhirui Zhang, Yunfei Lu, Xiaohui Yan, Duyu Tang, Dandan Tu, Bing Qin. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Yangfan Ye, Zekun Yuan, Xiachong Feng, Libo Qin 0001, Lei Huang 0021, Weitao Ma, Yichong Huang, Zhirui Zhang, Yunfei Lu, Duyu Tang, Dandan Tu, Bing Qin 0001 |
ACL (1) | 8 |
| 2025 | CLAIM: Mitigating Multilingual Object Hallucination in Large Vision-Language Models with Cross-Lingual Attention InterventionabstractLarge Vision-Language Models (LVLMs) have demonstrated impressive multimodal abilities but remain prone to multilingual object hallucination, with a higher likelihood of generating responses inconsistent with the visual input when utilizing queries in non-English languages compared to English. Most existing approaches to address these rely on pretraining or fine-tuning, which are resource-intensive. In this paper, inspired by observing the disparities in cross-modal attention patterns across languages, we propose Cross-Lingual Attention Intervention for Mitigating multilingual object hallucination (CLAIM) in LVLMs, a novel near training-free method by aligning attention patterns. CLAIM first identifies language-specific cross-modal attention heads, then estimates language shift vectors from English to the target language, and finally intervenes in the attention outputs during inference to facilitate cross-lingual visual perception capability alignment. Extensive experiments demonstrate that CLAIM achieves an average improvement of 13.56% (up to 30% in Spanish) on the POPE and 21.75% on the hallucination subsets of the MME benchmark across various languages. Further analysis reveals that multilingual attention divergence is most prominent in intermediate layers, highlighting their critical role in multilingual scenarios. Zekai Ye, Libo Qin 0001, Yichong Huang, Baohang Li, Kui Jiang, Yang Xiang 0003, Zhirui Zhang, Yunfei Lu, Duyu Tang, Dandan Tu, Bing Qin 0001 |
ACL (1) | 5 |
| 2025 | MCM-DPO: Multifaceted Cross-Modal Direct Preference Optimization for Alt-text Generation
Jinlan Fu, Shenzhen Huangfu, Hao Fei 0001, Yichong Huang, Xiaoyu Shen 0001, Xipeng Qiu, See-Kiong Ng |
ACM Multimedia | 4 |
| 2024 | Gradient Consistency-based Parameter Allocation for Multilingual Neural Machine TranslationabstractMultilingual neural machine translation handles the translation of multiple languages with one unified model. However, this joint-training paradigm incurs the notorious issue of parameter interference, where the model compromises with the language diversity to find a common solution. Recent research has explored avoiding this problem by selecting certain parameters for each language direction from the original model to form language-specific sub-networks. However, determining how many parameters to choose and which parameters to select is still a serious challenge. In this work, we propose an approach called CaPA (Consistency-based Parameter Allocation), which dynamically allocates parameters of appropriate scale to each language direction based on the consistency between the gradient of the individual language and the average gradient. Specifically, CaPA allocates more parameters to languages with higher gradient consistency as these languages tend to have a more positive impact on other languages. Furthermore, considering the varying levels of interference across different parts of the model, we propose an adaptive parameter allocation based on module-level gradient consistency. Experimental results show the correlation between gradient consistency and parameter interference, as well as the effectiveness of our proposed method. Wenshuai Huo, Yichong Huang, Chengpeng Fu, Hui Wang 0030, Bing Qin 0001 |
LREC/COLING | 3 |
| 2024 | Aligning Translation-Specific Understanding to General Understanding in Large Language ModelsabstractLarge Language models (LLMs) have exhibited remarkable abilities in understanding complex texts, offering a promising path towards human-like translation performance.However, this study reveals the misalignment between the translation-specific understanding and the general understanding inside LLMs.This understanding misalignment leads to LLMs mistakenly or literally translating some complicated concepts that they accurately comprehend in the general scenarios (e.g., QA).To align the translation-specific understanding to the general one, we propose a novel translation process, DUAT (Difficult words Understanding Aligned Translation), explicitly incorporating the general understanding on the complicated content incurring inconsistent understanding to guide the translation.Specifically, DUAT performs cross-lingual interpretation for the difficult-to-translate words and enhances the translation with the generated interpretations.Furthermore, we reframe the external tools to improve DUAT in detecting difficult words and generating helpful interpretations.We conduct experiments on the self-constructed benchmark Challenge-WMT 1 , consisting of samples that are prone to mistranslation.Human evaluation results on highresource and low-resource language pairs indicate that DUAT significantly facilitates the understanding alignment, which improves the translation quality (up to +3.85 COMET) and reduces the literality of the translation by -25% ∼ -51%. Yichong Huang, Baohang Li, Wenshuai Huo, Chengpeng Fu, Ting Liu 0001, Bing Qin 0001 |
EMNLP | 1 |
| 2024 | Ensemble Learning for Heterogeneous Large Language Models with Deep Parallel CollaborationabstractLarge language models (LLMs) exhibit complementary strengths in various tasks, motivating the research of LLM ensembling.
However, existing work focuses on training an extra reward model or fusion model to select or combine all candidate answers, posing a great challenge to the generalization on unseen data distributions.
Besides, prior methods use textual responses as communication media, ignoring the valuable information in the internal representations.
In this work, we propose a training-free ensemble framework \textsc{DeePEn}, fusing the informative probability distributions yielded by different LLMs at each decoding step.
Unfortunately, the vocabulary discrepancy between heterogeneous LLMs directly makes averaging the distributions unfeasible due to the token misalignment.
To address this challenge, \textsc{DeePEn} maps the probability distribution of each model from its own probability space to a universal \textit{relative space} based on the relative representation theory, and performs aggregation.
Next, we devise a search-based inverse transformation to transform the aggregated result back to the probability space of one of the ensembling LLMs (main model), in order to determine the next token.
We conduct extensive experiments on ensembles of different number of LLMs, ensembles of LLMs with different architectures, and ensembles between the LLM and the specialist model.
Experimental results show that (i) \textsc{DeePEn} achieves consistent improvements across six benchmarks covering subject examination, reasoning, and knowledge, (ii) a well-performing specialist model can benefit from a less effective LLM through distribution fusion, and (iii) \textsc{DeePEn} has complementary strengths with other ensemble methods such as voting. Yichong Huang, Baohang Li, Yang Xiang 0003, Hui Wang 0030, Ting Liu 0001, Bing Qin 0001 |
NeurIPS | 1 |
| 2023 | Towards Higher Pareto Frontier in Multilingual Machine TranslationabstractMultilingual neural machine translation has witnessed remarkable progress in recent years.However, the long-tailed distribution of multilingual corpora poses a challenge of Pareto optimization, i.e., optimizing for some languages may come at the cost of degrading the performance of others.Existing balancing training strategies are equivalent to a series of Pareto optimal solutions, which trade off on a Pareto frontier 1 .In this work, we propose a new training framework, Pareto Mutual Distillation (Pareto-MD), towards pushing the Pareto frontier outwards rather than making trade-offs.Specifically, Pareto-MD collaboratively trains two Pareto optimal solutions that favor different languages and allows them to learn from the strengths of each other via knowledge distillation.Furthermore, we introduce a novel strategy to enable stronger communication between Pareto optimal solutions and broaden the applicability of our approach.Experimental results on the widely-used WMT and TED datasets show that our method significantly pushes the Pareto frontier and outperforms baselines by up to +2.46 BLEU 2 . Yichong Huang, Xinwei Geng, Baohang Li, Bing Qin 0001 |
ACL (1) | 1 |
| 2022 | Unifying the Convergences in Multilingual Neural Machine TranslationabstractAlthough all-in-one-model multilingual neural machine translation (MNMT) has achieved remarkable progress, the convergence inconsistency in the joint training is ignored, i.e.,different language pairs reaching convergence in different epochs.This leads to the trained MNMT model over-fitting low-resource language translations while under-fitting highresource ones.In this paper, we propose a novel training strategy named LSSD (Language-Specific Self-Distillation), which can alleviate the convergence inconsistency and help MNMT models achieve the best performance on each language pair simultaneously.Specifically, LSSD picks up language-specific best checkpoints for each language pair to teach the current model on the fly.Furthermore, we systematically explore three sample-level manipulations of knowledge transferring.Experimental results on three datasets show that LSSD obtains consistent improvements towards all language pairs and achieves the state-of-the-art 1 . Yichong Huang, Xinwei Geng, Bing Qin 0001 |
EMNLP | 1 |