EDBT 2026 Demo / reviewers in the wild / expert
Yupu Liang
dblp:89/5863
· DBLP profile ↗
10ranked-venue papers
2as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Vision and language · 49% Machine translation · 24% Reinforcement learning · 13% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Bioinformatics and computational biology · 100% |
Topics — the 10 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language › vision-language model
multimodal large language model |
2.1 | 3 | 2026 | MT³: A Synergistic Multi-Task RL Framework for Specializing MLLMs in Text Image Machine Translation · ACL (1) 2026 SHIFT: Selected Helpful Informative Frame for Video-guided Machine Translation · EMNLP 2025 Single-to-mix Modality Alignment with Multimodal Large Language Model for Document Image Machine Translation · ACL (1) 2025 |
Machine learning › Reinforcement learning
multi-task reinforcement learning |
1.0 | 1 | 2026 | MT³: A Synergistic Multi-Task RL Framework for Specializing MLLMs in Text Image Machine Translation · ACL (1) 2026 |
Natural language and speech › Machine translation › multimodal machine translation
text image machine translation |
1.0 | 1 | 2026 | MT³: A Synergistic Multi-Task RL Framework for Specializing MLLMs in Text Image Machine Translation · ACL (1) 2026 |
Computer vision › Vision and language › multimodal understanding
multimodal document understanding |
0.9 | 1 | 2025 | Understand Layout and Translate Text: Unified Feature-Conductive End-to-End Document Image Translation · IEEE Trans. Pattern Anal. Mach. Intell. 2025 |
Natural language and speech › Machine translation › multimodal machine translation
video-guided machine translation |
0.9 | 1 | 2025 | SHIFT: Selected Helpful Informative Frame for Video-guided Machine Translation · EMNLP 2025 |
Computer vision › Vision and language › vision-language model
vision-language model alignment |
0.9 | 1 | 2025 | Single-to-mix Modality Alignment with Multimodal Large Language Model for Document Image Machine Translation · ACL (1) 2025 |
Bioinformatics and computational biology
negative sampling |
0.9 | 1 | 2025 | Topology-driven negative sampling enhances generalizability in protein-protein interaction prediction · Bioinform. 2025 |
Bioinformatics and computational biology
protein-protein interaction prediction |
0.9 | 1 | 2025 | Topology-driven negative sampling enhances generalizability in protein-protein interaction prediction · Bioinform. 2025 |
Bioinformatics and computational biology › protein analysis › protein bioinformatics
protein representation learning |
0.9 | 1 | 2025 | Topology-driven negative sampling enhances generalizability in protein-protein interaction prediction · Bioinform. 2025 |
Machine learning › Efficient and distributed learning
inference efficiency |
0.3 | 1 | 2025 | SHIFT: Selected Helpful Informative Frame for Video-guided Machine Translation · EMNLP 2025 |
Methods — techniques the papers use, named apart from their topics
multimodal large language model · 1.7reinforcement learning · 1.0multi-task learning · 1.0unsupervised pretraining · 0.9modality alignment · 0.9keyframe selection · 0.9graph machine learning · 0.9feature-conductive flow · 0.9end-to-end encoder-decoder · 0.9clustering · 0.9bridging mechanism · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MT³: A Synergistic Multi-Task RL Framework for Specializing MLLMs in Text Image Machine TranslationabstractZhaopeng Feng, Yupu Liang, Shaosheng Cao, Jiayuan Su, Jiahan Ren, Zhijie Zhou, Wenxuan Huang, Jian Wu, Zuozhu Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zhaopeng Feng, Yupu Liang, Shaosheng Cao, Jiayuan Su, Jiahan Ren, Wenxuan Huang 0001, Jian Wu 0001, Zuozhu Liu |
ACL (1) | 2 |
| 2025 | Single-to-mix Modality Alignment with Multimodal Large Language Model for Document Image Machine TranslationabstractYupu Liang, Yaping Zhang, Zhiyang Zhang, Yang Zhao, Lu Xiang, Chengqing Zong, Yu Zhou. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Yupu Liang, Yang Zhao 0007, Lu Xiang, Chengqing Zong, Yu Zhou 0001 |
ACL (1) | 1 |
| 2025 | From Chaotic OCR Words to Coherent Document: A Fine-to-Coarse Zoom-Out Network for Complex-Layout Document Image TranslationabstractDocument Image Translation (DIT) aims to translate documents in images from one language to another. It requires visual layouts and textual contents understanding, as well as document coherence capturing. However, current methods often rely on the quality of OCR output, which, particularly in complex-layout scenarios, frequently loses the crucial document coherence, leading to chaotic text. To overcome this problem, we introduce a novel end-to-end network, named Zoom-out DIT (ZoomDIT), inspired by human translation procedures. It jointly accomplishes the multi-level tasks including word positioning, sentence recognition & translation, and document organization, based on a fine-to-coarse zoom-out framework, to progressively realize “chaotic words to coherent document” and improve translation. We further contribute a new large-scale DIT dataset with multi-level fine-grained labels. Extensive experiments on public and our new dataset demonstrate significant improvements in translation quality towards complex-layout document images, offering a robust solution for reorganizing the chaotic OCR outputs to a coherent document translation. Yupu Liang, Lu Xiang, Yang Zhao 0007, Yu Zhou 0001, Chengqing Zong |
COLING | 3 |
| 2025 | SHIFT: Selected Helpful Informative Frame for Video-guided Machine TranslationabstractVideo-guided Machine Translation (VMT) aims to improve translation quality by integrating contextual information from paired short video clips.Mainstream VMT approaches typically incorporate multimodal information by uniformly sampling frames from the input videos.However, this paradigm frequently incurs significant computational overhead and introduces redundant multimodal content, which degrades both efficiency and translation quality.To tackle these challenges, we propose SHIFT (Selected Helpful Informative Frame for Translation).It is a lightweight, plug-andplay framework designed for VMT with Multimodal Large Language Models (MLLMs).SHIFT adaptively selects a single informative key frame when visual context is necessary; otherwise, it relies solely on textual input.This process is guided by a dedicated clustering module and a selector module.Experimental results demonstrate that SHIFT enhances the performance of MLLMs on the VMT task while simultaneously reducing computational cost, without sacrificing generalization ability. Boyu Guan, Chuang Han, Yupu Liang, Yang Zhao 0007, Chengqing Zong |
EMNLP | 4 |
| 2025 | ICDAR 2025 Competition on End-to-End Document Image Machine Translation Towards Complex Layouts
Yupu Liang, Lu Xiang, Yang Zhao 0007, Yu Zhou 0001, Chengqing Zong |
ICDAR (5) | 2 |
| 2025 | Boosting Document Image Translation via Layout-Aware Semantic Paragraph Clustering
Yupu Liang, Yunfei Lu, Dandan Tu, Chengqing Zong, Yu Zhou 0001 |
PRCV (7) | 4 |
| 2025 | Topology-driven negative sampling enhances generalizability in protein-protein interaction predictionabstractMOTIVATION: Unraveling the human interactome to uncover disease-specific patterns and discover drug targets hinges on accurate protein-protein interaction (PPI) predictions. However, challenges persist in machine learning (ML) models due to a scarcity of quality hard negative samples, shortcut learning, and limited generalizability to novel proteins. RESULTS: In this study, we introduce a novel approach for strategic sampling of protein-protein noninteractions (PPNIs) by leveraging higher-order network characteristics that capture the inherent complementarity-driven mechanisms of PPIs. Next, we introduce Unsupervised Pre-training of Node Attributes tuned for PPI (UPNA-PPI), a high throughput sequence-to-function ML pipeline, integrating unsupervised pre-training in protein representation learning with Topological PPNI (TPPNI) samples, capable of efficiently screening billions of interactions. By using our TPPNI in training the UPNA-PPI model, we improve PPI prediction generalizability and interpretability, particularly in identifying potential binding sites locations on amino acid sequences, strengthening the prioritization of screening assays and facilitating the transferability of ML predictions across protein families and homodimers. UPNA-PPI establishes the foundation for a fundamental negative sampling methodology in graph machine learning by integrating insights from network topology. AVAILABILITY AND IMPLEMENTATION: Code and UPNA-PPI predictions are freely available at https://github.com/alxndgb/UPNA-PPI. Babak Ravandi, Parham Haddadi, Naomi H. Philip, Mario Abdelmessih, William R. Mowrey, Piero Ricchiuto, Yupu Liang, Juan Carlos Mobarec, Tina Eliassi-Rad |
Bioinform. | 8 |
| 2025 | Understand Layout and Translate Text: Unified Feature-Conductive End-to-End Document Image TranslationabstractDocument Image Translation (DIT) aims to translate texts on document images from one language to another. It is a multi-modal task involving cooperation of text and layout. Current approaches either handle layout and translation as separate processes, risking accumulative errors, or use vanilla end-to-end encoder-decoder models to capture layout implicitly, often suffering inadequate layout incorporation. We argue that a favorable framework should explicitly engage layout-specific modules and properly organize them toward translation. For this, we first revisit two key layouts: the geometric layout reflecting word's spatial positions, and the logical layout depicting word's logical order. Then, a novel pipeline (understand layout $\rightarrow$→ translate text) is determined to prioritize layouts such that preceding layouts contribute to translation. Following this pipeline, we introduce Unified Document Image Translation (UniDIT), a comprehensive framework that unifies layout with translation in one network. It is devised to leverage each module's advantage, and provide an elaborate feature-conductive flow for module communication globally. A novel bridging mechanism is also introduced to adapt layout features conducive to translation. We further contribute DITransv2, a large-scale fine-grained benchmark that includes heterogeneous and complex document layouts. Extensive experiments on DITransv2 and additional established benchmarks demonstrate UniDIT outperforms previous state-of-the-arts in all aspects. Yupu Liang, Cong Ma 0002, Lu Xiang, Yang Zhao 0007, Yu Zhou 0001, Chengqing Zong |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Born a BabyNet with Hierarchical Parental Supervision for End-to-End Text Image Machine TranslationabstractText image machine translation (TIMT) aims at translating source language texts in images into another target language, which has been proven successful by bridging text image recognition encoder and text translation decoder. However, it is still an open question of how to incorporate fine-grained knowledge supervision to make it consistent between recognition and translation modules. In this paper, we propose a novel TIMT method named as BabyNet, which is optimized with hierarchical parental supervision to improve translation performance. Inspired by genetic recombination and variation in the field of genetics, the proposed BabyNet is inherited from the recognition and translation parent models with a variation module of which parameters can be updated when training on the TIMT task. Meanwhile, hierarchical and multi-granularity supervision from parent models is introduced to bridge the gap between inherited modules in BabyNet. Extensive experiments on both synthetic and real-world TIMT tests show that our proposed method significantly outperforms existing methods. Further analyses of various parent model combinations show the good generalization of our method. Cong Ma 0002, Yupu Liang, Yang Zhao 0007, Yu Zhou 0001, Chengqing Zong |
LREC/COLING | 4 |
| 2024 | Document Image Machine Translation with Dynamic Multi-pre-trained Models AssemblingabstractYupu Liang, Yaping Zhang, Cong Ma, Zhiyang Zhang, Yang Zhao, Lu Xiang, Chengqing Zong, Yu Zhou. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Yupu Liang, Cong Ma 0002, Yang Zhao 0007, Lu Xiang, Chengqing Zong, Yu Zhou 0001 |
NAACL-HLT | 1 |