EDBT 2026 Demo / reviewers in the wild / expert
Wenshuai Huo
dblp:362/8703
· DBLP profile ↗
5ranked-venue papers
3as first author
5since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Machine translation · 33% Language models and text generation · 31% Vision and language · 19% |
Topics — the 4 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language › image captioning
cross-lingual image captioning |
1.0 | 1 | 2026 | The Visual Prism: Refracting Images into Parallel Multilingual Descriptions with Structured Visual Guidance · AAAI 2026 |
Machine learning › Transfer learning and domain adaptation
fine-tuning |
0.9 | 1 | 2025 | Enhancing Non-English Capabilities of English-Centric Large Language Models Through Deep Supervision Fine-Tuning · AAAI 2025 |
Natural language and speech › Language models and text generation
multilingual language models |
0.9 | 1 | 2025 | Enhancing Non-English Capabilities of English-Centric Large Language Models Through Deep Supervision Fine-Tuning · AAAI 2025 |
Natural language and speech › Language models and text generation
large language model |
0.8 | 1 | 2024 | Aligning Translation-Specific Understanding to General Understanding in Large Language Models · EMNLP 2024 |
Methods — techniques the papers use, named apart from their topics
supervised fine-tuning · 1.0semantic graph guidance · 1.0reinforcement learning · 1.0logit supervision · 0.9feature supervision · 0.9deep supervision · 0.9external tools · 0.8cross-lingual interpretation · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | The Visual Prism: Refracting Images into Parallel Multilingual Descriptions with Structured Visual GuidanceabstractParallel corpora, as the foundation of machine translation, remain crucial even in the era of large language models (LLMs) for pre-training and fine-tuning. However, annotating parallel corpora is extremely costly, as it requires annotators to be proficient in multiple languages. To reduce this cost, prior work has explored image-pivoted corpus synthesis, generating multilingual captions for the same image as pseudo-parallel data. Unfortunately, these pseudo corpora suffer from the serious issue of multilingual focus divergence, i.e., the model attending to distinct aspects of the image when generating captions in different languages. To address this problem, we propose a method called PRISMS (Parallel Refracting ImageS into Multilingual descriptions with Structured visual guidance), which leverages semantic graphs as structured visual guidance to unify the focus of multilingual captions. To ensure adherence to this guidance, we introduce two key techniques: supervised fine-tuning using self-generated instructional data, and reinforcement learning with a reward signal based on semantic graph consistency. Experimental results on five languages show that our PRISMS significantly improves the image-pivot parallel corpora synthesis, enabling LLMs to achieve translation performance comparable to that of models trained on manually annotated corpora. Chengpeng Fu, Yichong Huang, Wenshuai Huo, Baohang Li, Yang Xiang 0003, Ting Liu 0001 |
AAAI | 4 |
| 2026 | English is not all you need: Rewarding better translation to inspire multilingual capability in LLMs
Wenshuai Huo, Yichong Huang, Chengpeng Fu, Hui Wang 0030, Bing Qin 0001 |
Neurocomputing | 1 |
| 2025 | Enhancing Non-English Capabilities of English-Centric Large Language Models Through Deep Supervision Fine-TuningabstractLarge language models (LLMs) have demonstrated significant progress in multilingual language understanding and generation. However, due to the imbalance in training data, their capabilities in non-English languages are limited. Recent studies revealed the English-pivot multilingual mechanism of LLMs, where LLMs implicitly convert non-English queries into English ones at the bottom layers and adopt English for thinking at the middle layers. However, due to the absence of explicit supervision for cross-lingual alignment in the intermediate layers of LLMs, the internal representations during these stages may become inaccurate. In this work, we introduce a deep supervision fine-tuning method (DFT) that incorporates additional supervision in the internal layers of the model to guide its workflow. Specifically, we introduce two training objectives on different layers of LLMs: one at the bottom layers to constrain the conversion of the target language into English, and another at the middle layers to constrain reasoning in English. To effectively achieve the guiding purpose, we designed two types of supervision signals: logits and feature, which represent a stricter constraint and a relatively more relaxed guidance. Our method guides the model to not only consider the final generated result when processing non-English inputs but also ensure the accuracy of internal representations. We conducted extensive experiments on typical English-centric large models, LLaMA-2 and Gemma-2, and the results on multiple multilingual datasets show that our method significantly outperforms traditional fine-tuning methods. Wenshuai Huo, Yichong Huang, Chengpeng Fu, Baohang Li, Yangfan Ye, Zhirui Zhang, Dandan Tu, Duyu Tang, Yunfei Lu, Hui Wang 0030, Bing Qin 0001 |
AAAI | 1 |
| 2024 | Gradient Consistency-based Parameter Allocation for Multilingual Neural Machine TranslationabstractMultilingual neural machine translation handles the translation of multiple languages with one unified model. However, this joint-training paradigm incurs the notorious issue of parameter interference, where the model compromises with the language diversity to find a common solution. Recent research has explored avoiding this problem by selecting certain parameters for each language direction from the original model to form language-specific sub-networks. However, determining how many parameters to choose and which parameters to select is still a serious challenge. In this work, we propose an approach called CaPA (Consistency-based Parameter Allocation), which dynamically allocates parameters of appropriate scale to each language direction based on the consistency between the gradient of the individual language and the average gradient. Specifically, CaPA allocates more parameters to languages with higher gradient consistency as these languages tend to have a more positive impact on other languages. Furthermore, considering the varying levels of interference across different parts of the model, we propose an adaptive parameter allocation based on module-level gradient consistency. Experimental results show the correlation between gradient consistency and parameter interference, as well as the effectiveness of our proposed method. Wenshuai Huo, Yichong Huang, Chengpeng Fu, Hui Wang 0030, Bing Qin 0001 |
LREC/COLING | 1 |
| 2024 | Aligning Translation-Specific Understanding to General Understanding in Large Language ModelsabstractLarge Language models (LLMs) have exhibited remarkable abilities in understanding complex texts, offering a promising path towards human-like translation performance.However, this study reveals the misalignment between the translation-specific understanding and the general understanding inside LLMs.This understanding misalignment leads to LLMs mistakenly or literally translating some complicated concepts that they accurately comprehend in the general scenarios (e.g., QA).To align the translation-specific understanding to the general one, we propose a novel translation process, DUAT (Difficult words Understanding Aligned Translation), explicitly incorporating the general understanding on the complicated content incurring inconsistent understanding to guide the translation.Specifically, DUAT performs cross-lingual interpretation for the difficult-to-translate words and enhances the translation with the generated interpretations.Furthermore, we reframe the external tools to improve DUAT in detecting difficult words and generating helpful interpretations.We conduct experiments on the self-constructed benchmark Challenge-WMT 1 , consisting of samples that are prone to mistranslation.Human evaluation results on highresource and low-resource language pairs indicate that DUAT significantly facilitates the understanding alignment, which improves the translation quality (up to +3.85 COMET) and reduces the literality of the translation by -25% ∼ -51%. Yichong Huang, Baohang Li, Wenshuai Huo, Chengpeng Fu, Ting Liu 0001, Bing Qin 0001 |
EMNLP | 4 |