VLDB 2026 Research / reviewers in the wild / expert
Linjuan Wu
dblp:262/2608
· DBLP profile ↗
10ranked-venue papers
4as first author
10since 2021 · last 2026
0000-0002-2168-3045ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 4 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
9 papers |
Language models and text generation · 47% Transfer learning and domain adaptation · 21% Representation and self-supervised learning · 8% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computing education · 100% | |
| Computer graphics and multimedia
1 paper |
Visual content generation and editing · 100% |
Topics — the 20 heaviest of 23, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Transfer learning and domain adaptation
cross-lingual transfer |
3.3 | 5 | 2026 | Enhancing LLM Language Adaption through Cross-lingual In-Context Pre-training · EMNLP 2025 From English to Second Language Mastery: Enhancing LLMs with Cross-Lingual Continued Instruction Tuning · ACL (1) 2025 Struct-XLM: A Structure Discovery Multilingual Language Model for Enhancing Cross-lingual Transfer through Reinforcement Learning · EMNLP 2023 |
Natural language and speech › Language models and text generation
multilingual language models |
1.9 | 2 | 2026 | A Data-Efficient Path to Multilingual LLMs: Language Expansion via Post-training PARAMΔ Integration into Upcycled MoE · ACL (1) 2026 Enhancing LLM Language Adaption through Cross-lingual In-Context Pre-training · EMNLP 2025 |
Natural language and speech › Language models and text generation › multilingual language models
language expansion |
1.0 | 1 | 2026 | A Data-Efficient Path to Multilingual LLMs: Language Expansion via Post-training PARAMΔ Integration into Upcycled MoE · ACL (1) 2026 |
Natural language and speech › Language models and text generation
large language model evaluation |
1.0 | 1 | 2026 | AutoTaskEval: Towards Domain-Specific and Fine-Grained Evaluation for LLMs · ACL (1) 2026 |
Machine learning › Efficient and distributed learning › model reuse
model upcycling |
1.0 | 1 | 2026 | A Data-Efficient Path to Multilingual LLMs: Language Expansion via Post-training PARAMΔ Integration into Upcycled MoE · ACL (1) 2026 |
Natural language and speech › Language models and text generation › large language model training › language model pretraining
in-context pretraining |
0.9 | 1 | 2025 | Enhancing LLM Language Adaption through Cross-lingual In-Context Pre-training · EMNLP 2025 |
Natural language and speech › Language models and text generation
instruction tuning |
0.9 | 1 | 2025 | From English to Second Language Mastery: Enhancing LLMs with Cross-Lingual Continued Instruction Tuning · ACL (1) 2025 |
Computer vision › Vision and language › vision-language model
multimodal large language model |
0.9 | 1 | 2025 | SVGenius: Benchmarking LLMs in SVG Understanding, Editing and Generation · ACM Multimedia 2025 |
Natural language and speech › Language models and text generation › LLM agents
tool use |
0.9 | 1 | 2025 | AskToAct: Enhancing LLMs Tool Use via Self-Correcting Clarification · EMNLP 2025 |
Computing education
large language model evaluation |
0.9 | 1 | 2025 | SVGenius: Benchmarking LLMs in SVG Understanding, Editing and Generation · ACM Multimedia 2025 |
Visual content generation and editing
vector graphics generation |
0.9 | 1 | 2025 | SVGenius: Benchmarking LLMs in SVG Understanding, Editing and Generation · ACM Multimedia 2025 |
Natural language and speech › Language models and text generation
self-reflection |
0.8 | 1 | 2024 | Self-Contrast: Better Reflection Through Inconsistent Solving Perspectives · ACL (1) 2024 |
Machine learning › Representation and self-supervised learning › representation matching › feature alignment › embedding alignment
cross-lingual alignment |
0.7 | 1 | 2023 | Struct-XLM: A Structure Discovery Multilingual Language Model for Enhancing Cross-lingual Transfer through Reinforcement Learning · EMNLP 2023 |
Natural language and speech › Language models and text generation › multilingual language models
multilingual pretrained language model |
0.7 | 1 | 2023 | Struct-XLM: A Structure Discovery Multilingual Language Model for Enhancing Cross-lingual Transfer through Reinforcement Learning · EMNLP 2023 |
Natural language and speech › Information extraction and text analysis › syntactic parsing
syntactic structure induction |
0.7 | 1 | 2023 | Struct-XLM: A Structure Discovery Multilingual Language Model for Enhancing Cross-lingual Transfer through Reinforcement Learning · EMNLP 2023 |
Natural language and speech › Question answering and dialogue systems › machine reading comprehension
cross-lingual machine reading comprehension |
0.6 | 1 | 2022 | Learning Disentangled Semantic Representations for Zero-Shot Cross-Lingual Transfer in Multilingual Machine Reading Comprehension · ACL (1) 2022 |
Machine learning › Representation and self-supervised learning › representation learning
disentangled representation learning |
0.6 | 1 | 2022 | Learning Disentangled Semantic Representations for Zero-Shot Cross-Lingual Transfer in Multilingual Machine Reading Comprehension · ACL (1) 2022 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
semantic representation |
0.6 | 1 | 2022 | Learning Disentangled Semantic Representations for Zero-Shot Cross-Lingual Transfer in Multilingual Machine Reading Comprehension · ACL (1) 2022 |
Machine learning › Transfer learning and domain adaptation › cross-lingual transfer
zero-shot cross-lingual transfer |
0.6 | 1 | 2022 | Learning Disentangled Semantic Representations for Zero-Shot Cross-Lingual Transfer in Multilingual Machine Reading Comprehension · ACL (1) 2022 |
Machine learning › Representation and self-supervised learning
pre-training |
0.3 | 1 | 2025 | Enhancing LLM Language Adaption through Cross-lingual In-Context Pre-training · EMNLP 2025 |
Methods — techniques the papers use, named apart from their topics
post-training parameter integration · 1.0mixture of experts · 1.0benchmark construction · 1.0supervised fine-tuning · 0.9semantic retrieval · 0.9self-paced learning · 0.9next-word prediction · 0.9in-context learning · 0.9LLM-as-a-judge · 0.9self-contrast · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AutoTaskEval: Towards Domain-Specific and Fine-Grained Evaluation for LLMsabstractQingqing Lyu, Linjuan Wu, Yongliang Shen, Hengwei Liu, Hao Li, Shengpei Jiang, Yin Zhang, Weiming Lu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Qingqing Lyu, Linjuan Wu, Yongliang Shen 0001, Hengwei Liu, Shengpei Jiang, Yin Zhang 0006, Weiming Lu 0001 |
ACL (1) | 2 |
| 2026 | A Data-Efficient Path to Multilingual LLMs: Language Expansion via Post-training PARAMΔ Integration into Upcycled MoEabstractHao Zhou, Tianhao Li, Zhijun Wang, Shuaijie She, Linjuan Wu, Hao-Ran Wei, Baosong Yang, Jiajun Chen, Shujian Huang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Hao Zhou 0012, Shuaijie She, Linjuan Wu, Baosong Yang, Jiajun Chen 0001, Shujian Huang |
ACL (1) | 5 |
| 2025 | From English to Second Language Mastery: Enhancing LLMs with Cross-Lingual Continued Instruction TuningabstractSupervised Fine-Tuning (SFT) with translated instruction data effectively adapts Large Language Models (LLMs) from English to non-English languages.We introduce Cross-Lingual Continued Instruction Tuning (X-CIT), which fully leverages translation-based parallel instruction data to enhance cross-lingual adaptability.X-CIT emulates the human process of second language acquisition and is guided by Chomsky's Principles and Parameters Theory.It first fine-tunes the LLM on English instruction data to establish foundational capabilities (i.e.Principles), then continues with target language translation and customized chatinstruction data to adjust "parameters" specific to the target language.This chat-instruction data captures alignment information in translated parallel data, guiding the model to initially think and respond in its native language before transitioning to the target language.To further mimic human learning progression, we incorporate Self-Paced Learning (SPL) during continued training, allowing the model to advance from simple to complex tasks.Implemented on Llama-2-7B across five languages, X-CIT was evaluated against three objective benchmarks and an LLM-as-a-judge benchmark, improving the strongest baseline by an average of 1.97% and 8.2% in these two benchmarks, respectively. Linjuan Wu, Baosong Yang, Weiming Lu 0001 |
ACL (1) | 1 |
| 2025 | Enhancing LLM Language Adaption through Cross-lingual In-Context Pre-trainingabstractLarge language models (LLMs) exhibit remarkable multilingual capabilities despite Englishdominated pre-training, attributed to crosslingual mechanisms during pre-training.Existing methods for enhancing cross-lingual transfer remain constrained by parallel resources, suffering from limited linguistic and domain coverage.We propose Cross-lingual In-context Pre-training (CrossIC-PT), a simple and scalable approach that enhances cross-lingual transfer by leveraging semantically related bilingual texts via simple next-word prediction.We construct CrossIC-PT samples by interleaving semantic-related bilingual Wikipedia documents into a single context window.To access window size constraints, we implement a systematic segmentation policy to split long bilingual document pairs into chunks while adjusting the sliding window mechanism to preserve contextual coherence.We further extend data availability through a semantic retrieval framework to construct CrossIC-PT samples from web-crawled corpus.Experimental results demonstrate that CrossIC-PT improves multilingual performance on three models (Llama-3.1-8B,Qwen2.5-7B, and Qwen2.5-1.5B)across six target languages, yielding performance gains of 3.79%, 3.99%, and 1.95%, respectively, with additional improvements after data augmentation. Linjuan Wu, Baosong Yang, Fei Huang 0002, Weiming Lu 0001 |
EMNLP | 1 |
| 2025 | AskToAct: Enhancing LLMs Tool Use via Self-Correcting ClarificationabstractXuan Zhang, Yongliang Shen, Zhe Zheng, Linjuan Wu, Wenqi Zhang, Yuchen Yan, Qiuying Peng, Jun Wang, Weiming Lu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Yongliang Shen 0001, Linjuan Wu, Wenqi Zhang 0001, Qiuying Peng, Weiming Lu 0001 |
EMNLP | 4 |
| 2025 | SVGenius: Benchmarking LLMs in SVG Understanding, Editing and GenerationabstractLarge Language Models (LLMs) and Multimodal LLMs have shown promising capabilities for SVG processing, yet existing benchmarks suffer from limited real-world coverage, lack of complexity stratification, and fragmented evaluation paradigms. We introduce SVGenius, a comprehensive benchmark comprising 2,377 queries across three progressive dimensions: understanding, editing, and generation. Built on real-world data from 24 application domains with systematic complexity stratification, SVGenius evaluates models through 8 task categories and 18 metrics. We assess 22 mainstream models spanning different scales, architectures, training paradigms, and accessibility levels. Our analysis reveals that while proprietary models significantly outperform open-source counterparts, all models exhibit systematic performance degradation with increasing complexity, indicating fundamental limitations in current approaches; however, reasoning-enhanced training proves more effective than pure scaling for overcoming these limitations, though style transfer remains the most challenging capability across all model types. SVGenius establishes the first systematic evaluation framework for SVG processing, providing crucial insights for developing more capable vector graphics models and advancing automated graphic design applications. Appendix and supplementary materials (including all data and code) are available at https://zju-real.github.io/SVGenius. Haolei Xu, Fei Tang 0005, Linjuan Wu, Wenqi Zhang 0001, Guiyang Hou, Yongliang Shen 0001, Weiming Lu 0001, Yueting Zhuang |
ACM Multimedia | 8 |
| 2024 | Self-Contrast: Better Reflection Through Inconsistent Solving PerspectivesabstractWenqi Zhang, Yongliang Shen, Linjuan Wu, Qiuying Peng, Jun Wang, Yueting Zhuang, Weiming Lu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Wenqi Zhang 0001, Yongliang Shen 0001, Linjuan Wu, Qiuying Peng, Yueting Zhuang, Weiming Lu 0001 |
ACL (1) | 3 |
| 2023 | Struct-XLM: A Structure Discovery Multilingual Language Model for Enhancing Cross-lingual Transfer through Reinforcement LearningabstractCross-lingual transfer learning heavily relies on well-aligned cross-lingual representations.The syntactic structure is recognized as beneficial for cross-lingual transfer, but limited researches utilize it for aligning representation in multilingual pre-trained language models (PLMs).Additionally, existing methods require syntactic labels that are difficult to obtain and of poor quality for low-resource languages.To address this gap, we propose Struct-XLM, a novel multilingual language model that leverages reinforcement learning (RL) to autonomously discover universal syntactic structures for improving the cross-lingual representation alignment of PLM.Struct-XLM integrates a policy network (PNet) and a translation ranking task.The PNet is designed to discover structural information and integrate it into the last layer of the PLM through the structural multi-head attention module to obtain structural representation.The translation ranking task obtains a delayed reward based on the structural representation to optimize the PNet while improving the alignment of cross-lingual representation.Experiments show the effectiveness of the proposed approach for enhancing cross-lingual transfer of multilingual PLM on the XTREME benchmark 1 . Linjuan Wu, Weiming Lu 0001 |
EMNLP | 1 |
| 2022 | Learning Disentangled Semantic Representations for Zero-Shot Cross-Lingual Transfer in Multilingual Machine Reading ComprehensionabstractLinjuan Wu, Shaojuan Wu, Xiaowang Zhang, Deyi Xiong, Shizhan Chen, Zhiqiang Zhuang, Zhiyong Feng. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Linjuan Wu, Shaojuan Wu, Xiaowang Zhang, Deyi Xiong, Shizhan Chen, Zhiqiang Zhuang, Zhiyong Feng 0002 |
ACL (1) | 1 |
| 2021 | Modeling Global Semantics for Question Answering over Knowledge BasesabstractQuery graph as a junction of semantic parsing in question answering over knowledge bases (KBQA) connects questions and logical queries. Though query graph consists of rich information such as structure, relation, etc., the current KBQA's models mainly utilize limited relation information in a naive way. It is not easy to learn the representation of a query graph with that information due to the heterogeneity of the query graph and intricate correlation of relations. In this paper, we propose a Global Semantic-based Message Passing (GSMP) model to model the global semantics of a query graph from its structure and relation information. In GSMP, we present a recurrent-based relational graph convolutional network (RGCN) to capture heterogeneous query graphs where the recurrent unit improves the capability of RGCN in processing small-scale query graphs. Moreover, we present a contextual-based method to remove ambiguity caused by intricate correlations where the contextual adjacency of relations optimizes relation representation. Finally, we present a nonlinear gate-based encoder to learning the representation of questions' syntactic tree, as the structure information of questions, for better matching the global semantics of query graphs. Experiments evaluated on benchmarks show that our model outperforms off-the-shelf models. Peiyun Wu, Yunjie Wu, Linjuan Wu, Xiaowang Zhang, Zhiyong Feng 0002 |
IJCNN | 3 |