Yupu Liang

dblp:89/5863 · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Vision and language · 49% Machine translation · 24% Reinforcement learning · 13%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%

Topics — the 10 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language › vision-language model
multimodal large language model
2.132026
MT³: A Synergistic Multi-Task RL Framework for Specializing MLLMs in Text Image Machine Translation · ACL (1) 2026
SHIFT: Selected Helpful Informative Frame for Video-guided Machine Translation · EMNLP 2025
Single-to-mix Modality Alignment with Multimodal Large Language Model for Document Image Machine Translation · ACL (1) 2025
Machine learning › Reinforcement learning
multi-task reinforcement learning
1.012026
MT³: A Synergistic Multi-Task RL Framework for Specializing MLLMs in Text Image Machine Translation · ACL (1) 2026
Natural language and speech › Machine translation › multimodal machine translation
text image machine translation
1.012026
MT³: A Synergistic Multi-Task RL Framework for Specializing MLLMs in Text Image Machine Translation · ACL (1) 2026
Computer vision › Vision and language › multimodal understanding
multimodal document understanding
0.912025
Understand Layout and Translate Text: Unified Feature-Conductive End-to-End Document Image Translation · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Natural language and speech › Machine translation › multimodal machine translation
video-guided machine translation
0.912025
SHIFT: Selected Helpful Informative Frame for Video-guided Machine Translation · EMNLP 2025
Computer vision › Vision and language › vision-language model
vision-language model alignment
0.912025
Single-to-mix Modality Alignment with Multimodal Large Language Model for Document Image Machine Translation · ACL (1) 2025
Bioinformatics and computational biology
negative sampling
0.912025
Topology-driven negative sampling enhances generalizability in protein-protein interaction prediction · Bioinform. 2025
Bioinformatics and computational biology
protein-protein interaction prediction
0.912025
Topology-driven negative sampling enhances generalizability in protein-protein interaction prediction · Bioinform. 2025
Bioinformatics and computational biology › protein analysis › protein bioinformatics
protein representation learning
0.912025
Topology-driven negative sampling enhances generalizability in protein-protein interaction prediction · Bioinform. 2025
Machine learning › Efficient and distributed learning
inference efficiency
0.312025
SHIFT: Selected Helpful Informative Frame for Video-guided Machine Translation · EMNLP 2025

Methods — techniques the papers use, named apart from their topics

multimodal large language model · 1.7reinforcement learning · 1.0multi-task learning · 1.0unsupervised pretraining · 0.9modality alignment · 0.9keyframe selection · 0.9graph machine learning · 0.9feature-conductive flow · 0.9end-to-end encoder-decoder · 0.9clustering · 0.9bridging mechanism · 0.9
YearPublicationVenuePosition
2026 MT³: A Synergistic Multi-Task RL Framework for Specializing MLLMs in Text Image Machine Translation
abstract
Zhaopeng Feng, Yupu Liang, Shaosheng Cao, Jiayuan Su, Jiahan Ren, Zhijie Zhou, Wenxuan Huang, Jian Wu, Zuozhu Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zhaopeng Feng, Yupu Liang, Shaosheng Cao, Jiayuan Su, Jiahan Ren, Wenxuan Huang 0001, Jian Wu 0001, Zuozhu Liu
ACL (1)2
2025 Single-to-mix Modality Alignment with Multimodal Large Language Model for Document Image Machine Translation
abstract
Yupu Liang, Yaping Zhang, Zhiyang Zhang, Yang Zhao, Lu Xiang, Chengqing Zong, Yu Zhou. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Yupu Liang, Yang Zhao 0007, Lu Xiang, Chengqing Zong, Yu Zhou 0001
ACL (1)1
2025 From Chaotic OCR Words to Coherent Document: A Fine-to-Coarse Zoom-Out Network for Complex-Layout Document Image Translation
abstract
Document Image Translation (DIT) aims to translate documents in images from one language to another. It requires visual layouts and textual contents understanding, as well as document coherence capturing. However, current methods often rely on the quality of OCR output, which, particularly in complex-layout scenarios, frequently loses the crucial document coherence, leading to chaotic text. To overcome this problem, we introduce a novel end-to-end network, named Zoom-out DIT (ZoomDIT), inspired by human translation procedures. It jointly accomplishes the multi-level tasks including word positioning, sentence recognition & translation, and document organization, based on a fine-to-coarse zoom-out framework, to progressively realize “chaotic words to coherent document” and improve translation. We further contribute a new large-scale DIT dataset with multi-level fine-grained labels. Extensive experiments on public and our new dataset demonstrate significant improvements in translation quality towards complex-layout document images, offering a robust solution for reorganizing the chaotic OCR outputs to a coherent document translation.
Yupu Liang, Lu Xiang, Yang Zhao 0007, Yu Zhou 0001, Chengqing Zong
COLING3
2025 SHIFT: Selected Helpful Informative Frame for Video-guided Machine Translation
abstract
Video-guided Machine Translation (VMT) aims to improve translation quality by integrating contextual information from paired short video clips.Mainstream VMT approaches typically incorporate multimodal information by uniformly sampling frames from the input videos.However, this paradigm frequently incurs significant computational overhead and introduces redundant multimodal content, which degrades both efficiency and translation quality.To tackle these challenges, we propose SHIFT (Selected Helpful Informative Frame for Translation).It is a lightweight, plug-andplay framework designed for VMT with Multimodal Large Language Models (MLLMs).SHIFT adaptively selects a single informative key frame when visual context is necessary; otherwise, it relies solely on textual input.This process is guided by a dedicated clustering module and a selector module.Experimental results demonstrate that SHIFT enhances the performance of MLLMs on the VMT task while simultaneously reducing computational cost, without sacrificing generalization ability.
Boyu Guan, Chuang Han, Yupu Liang, Yang Zhao 0007, Chengqing Zong
EMNLP4
2025 ICDAR 2025 Competition on End-to-End Document Image Machine Translation Towards Complex Layouts
Yupu Liang, Lu Xiang, Yang Zhao 0007, Yu Zhou 0001, Chengqing Zong
ICDAR (5)2
2025 Boosting Document Image Translation via Layout-Aware Semantic Paragraph Clustering
Yupu Liang, Yunfei Lu, Dandan Tu, Chengqing Zong, Yu Zhou 0001
PRCV (7)4
2025 Topology-driven negative sampling enhances generalizability in protein-protein interaction prediction
abstract
MOTIVATION: Unraveling the human interactome to uncover disease-specific patterns and discover drug targets hinges on accurate protein-protein interaction (PPI) predictions. However, challenges persist in machine learning (ML) models due to a scarcity of quality hard negative samples, shortcut learning, and limited generalizability to novel proteins. RESULTS: In this study, we introduce a novel approach for strategic sampling of protein-protein noninteractions (PPNIs) by leveraging higher-order network characteristics that capture the inherent complementarity-driven mechanisms of PPIs. Next, we introduce Unsupervised Pre-training of Node Attributes tuned for PPI (UPNA-PPI), a high throughput sequence-to-function ML pipeline, integrating unsupervised pre-training in protein representation learning with Topological PPNI (TPPNI) samples, capable of efficiently screening billions of interactions. By using our TPPNI in training the UPNA-PPI model, we improve PPI prediction generalizability and interpretability, particularly in identifying potential binding sites locations on amino acid sequences, strengthening the prioritization of screening assays and facilitating the transferability of ML predictions across protein families and homodimers. UPNA-PPI establishes the foundation for a fundamental negative sampling methodology in graph machine learning by integrating insights from network topology. AVAILABILITY AND IMPLEMENTATION: Code and UPNA-PPI predictions are freely available at https://github.com/alxndgb/UPNA-PPI.
Babak Ravandi, Parham Haddadi, Naomi H. Philip, Mario Abdelmessih, William R. Mowrey, Piero Ricchiuto, Yupu Liang, Juan Carlos Mobarec, Tina Eliassi-Rad
Bioinform.8
2025 Understand Layout and Translate Text: Unified Feature-Conductive End-to-End Document Image Translation
abstract
Document Image Translation (DIT) aims to translate texts on document images from one language to another. It is a multi-modal task involving cooperation of text and layout. Current approaches either handle layout and translation as separate processes, risking accumulative errors, or use vanilla end-to-end encoder-decoder models to capture layout implicitly, often suffering inadequate layout incorporation. We argue that a favorable framework should explicitly engage layout-specific modules and properly organize them toward translation. For this, we first revisit two key layouts: the geometric layout reflecting word's spatial positions, and the logical layout depicting word's logical order. Then, a novel pipeline (understand layout $\rightarrow$→ translate text) is determined to prioritize layouts such that preceding layouts contribute to translation. Following this pipeline, we introduce Unified Document Image Translation (UniDIT), a comprehensive framework that unifies layout with translation in one network. It is devised to leverage each module's advantage, and provide an elaborate feature-conductive flow for module communication globally. A novel bridging mechanism is also introduced to adapt layout features conducive to translation. We further contribute DITransv2, a large-scale fine-grained benchmark that includes heterogeneous and complex document layouts. Extensive experiments on DITransv2 and additional established benchmarks demonstrate UniDIT outperforms previous state-of-the-arts in all aspects.
Yupu Liang, Cong Ma 0002, Lu Xiang, Yang Zhao 0007, Yu Zhou 0001, Chengqing Zong
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 Born a BabyNet with Hierarchical Parental Supervision for End-to-End Text Image Machine Translation
abstract
Text image machine translation (TIMT) aims at translating source language texts in images into another target language, which has been proven successful by bridging text image recognition encoder and text translation decoder. However, it is still an open question of how to incorporate fine-grained knowledge supervision to make it consistent between recognition and translation modules. In this paper, we propose a novel TIMT method named as BabyNet, which is optimized with hierarchical parental supervision to improve translation performance. Inspired by genetic recombination and variation in the field of genetics, the proposed BabyNet is inherited from the recognition and translation parent models with a variation module of which parameters can be updated when training on the TIMT task. Meanwhile, hierarchical and multi-granularity supervision from parent models is introduced to bridge the gap between inherited modules in BabyNet. Extensive experiments on both synthetic and real-world TIMT tests show that our proposed method significantly outperforms existing methods. Further analyses of various parent model combinations show the good generalization of our method.
Cong Ma 0002, Yupu Liang, Yang Zhao 0007, Yu Zhou 0001, Chengqing Zong
LREC/COLING4
2024 Document Image Machine Translation with Dynamic Multi-pre-trained Models Assembling
abstract
Yupu Liang, Yaping Zhang, Cong Ma, Zhiyang Zhang, Yang Zhao, Lu Xiang, Chengqing Zong, Yu Zhou. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Yupu Liang, Cong Ma 0002, Yang Zhao 0007, Lu Xiang, Chengqing Zong, Yu Zhou 0001
NAACL-HLT1