Linjuan Wu

dblp:262/2608 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
10since 2021 · last 2026
0000-0002-2168-3045ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 4 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
9 papers
Language models and text generation · 47% Transfer learning and domain adaptation · 21% Representation and self-supervised learning · 8%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computing education · 100%
Computer graphics and multimedia
1 paper
Visual content generation and editing · 100%

Topics — the 20 heaviest of 23, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Transfer learning and domain adaptation
cross-lingual transfer
3.352026
Enhancing LLM Language Adaption through Cross-lingual In-Context Pre-training · EMNLP 2025
From English to Second Language Mastery: Enhancing LLMs with Cross-Lingual Continued Instruction Tuning · ACL (1) 2025
Struct-XLM: A Structure Discovery Multilingual Language Model for Enhancing Cross-lingual Transfer through Reinforcement Learning · EMNLP 2023
Natural language and speech › Language models and text generation
multilingual language models
1.922026
A Data-Efficient Path to Multilingual LLMs: Language Expansion via Post-training PARAMΔ Integration into Upcycled MoE · ACL (1) 2026
Enhancing LLM Language Adaption through Cross-lingual In-Context Pre-training · EMNLP 2025
Natural language and speech › Language models and text generation › multilingual language models
language expansion
1.012026
A Data-Efficient Path to Multilingual LLMs: Language Expansion via Post-training PARAMΔ Integration into Upcycled MoE · ACL (1) 2026
Natural language and speech › Language models and text generation
large language model evaluation
1.012026
AutoTaskEval: Towards Domain-Specific and Fine-Grained Evaluation for LLMs · ACL (1) 2026
Machine learning › Efficient and distributed learning › model reuse
model upcycling
1.012026
A Data-Efficient Path to Multilingual LLMs: Language Expansion via Post-training PARAMΔ Integration into Upcycled MoE · ACL (1) 2026
Natural language and speech › Language models and text generation › large language model training › language model pretraining
in-context pretraining
0.912025
Enhancing LLM Language Adaption through Cross-lingual In-Context Pre-training · EMNLP 2025
Natural language and speech › Language models and text generation
instruction tuning
0.912025
From English to Second Language Mastery: Enhancing LLMs with Cross-Lingual Continued Instruction Tuning · ACL (1) 2025
Computer vision › Vision and language › vision-language model
multimodal large language model
0.912025
SVGenius: Benchmarking LLMs in SVG Understanding, Editing and Generation · ACM Multimedia 2025
Natural language and speech › Language models and text generation › LLM agents
tool use
0.912025
AskToAct: Enhancing LLMs Tool Use via Self-Correcting Clarification · EMNLP 2025
Computing education
large language model evaluation
0.912025
SVGenius: Benchmarking LLMs in SVG Understanding, Editing and Generation · ACM Multimedia 2025
Visual content generation and editing
vector graphics generation
0.912025
SVGenius: Benchmarking LLMs in SVG Understanding, Editing and Generation · ACM Multimedia 2025
Natural language and speech › Language models and text generation
self-reflection
0.812024
Self-Contrast: Better Reflection Through Inconsistent Solving Perspectives · ACL (1) 2024
Machine learning › Representation and self-supervised learning › representation matching › feature alignment › embedding alignment
cross-lingual alignment
0.712023
Struct-XLM: A Structure Discovery Multilingual Language Model for Enhancing Cross-lingual Transfer through Reinforcement Learning · EMNLP 2023
Natural language and speech › Language models and text generation › multilingual language models
multilingual pretrained language model
0.712023
Struct-XLM: A Structure Discovery Multilingual Language Model for Enhancing Cross-lingual Transfer through Reinforcement Learning · EMNLP 2023
Natural language and speech › Information extraction and text analysis › syntactic parsing
syntactic structure induction
0.712023
Struct-XLM: A Structure Discovery Multilingual Language Model for Enhancing Cross-lingual Transfer through Reinforcement Learning · EMNLP 2023
Natural language and speech › Question answering and dialogue systems › machine reading comprehension
cross-lingual machine reading comprehension
0.612022
Learning Disentangled Semantic Representations for Zero-Shot Cross-Lingual Transfer in Multilingual Machine Reading Comprehension · ACL (1) 2022
Machine learning › Representation and self-supervised learning › representation learning
disentangled representation learning
0.612022
Learning Disentangled Semantic Representations for Zero-Shot Cross-Lingual Transfer in Multilingual Machine Reading Comprehension · ACL (1) 2022
Knowledge, reasoning and agents › Knowledge representation and reasoning
semantic representation
0.612022
Learning Disentangled Semantic Representations for Zero-Shot Cross-Lingual Transfer in Multilingual Machine Reading Comprehension · ACL (1) 2022
Machine learning › Transfer learning and domain adaptation › cross-lingual transfer
zero-shot cross-lingual transfer
0.612022
Learning Disentangled Semantic Representations for Zero-Shot Cross-Lingual Transfer in Multilingual Machine Reading Comprehension · ACL (1) 2022
Machine learning › Representation and self-supervised learning
pre-training
0.312025
Enhancing LLM Language Adaption through Cross-lingual In-Context Pre-training · EMNLP 2025

Methods — techniques the papers use, named apart from their topics

post-training parameter integration · 1.0mixture of experts · 1.0benchmark construction · 1.0supervised fine-tuning · 0.9semantic retrieval · 0.9self-paced learning · 0.9next-word prediction · 0.9in-context learning · 0.9LLM-as-a-judge · 0.9self-contrast · 0.8
YearPublicationVenuePosition
2026 AutoTaskEval: Towards Domain-Specific and Fine-Grained Evaluation for LLMs
abstract
Qingqing Lyu, Linjuan Wu, Yongliang Shen, Hengwei Liu, Hao Li, Shengpei Jiang, Yin Zhang, Weiming Lu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Qingqing Lyu, Linjuan Wu, Yongliang Shen 0001, Hengwei Liu, Shengpei Jiang, Yin Zhang 0006, Weiming Lu 0001
ACL (1)2
2026 A Data-Efficient Path to Multilingual LLMs: Language Expansion via Post-training PARAMΔ Integration into Upcycled MoE
abstract
Hao Zhou, Tianhao Li, Zhijun Wang, Shuaijie She, Linjuan Wu, Hao-Ran Wei, Baosong Yang, Jiajun Chen, Shujian Huang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Hao Zhou 0012, Shuaijie She, Linjuan Wu, Baosong Yang, Jiajun Chen 0001, Shujian Huang
ACL (1)5
2025 From English to Second Language Mastery: Enhancing LLMs with Cross-Lingual Continued Instruction Tuning
abstract
Supervised Fine-Tuning (SFT) with translated instruction data effectively adapts Large Language Models (LLMs) from English to non-English languages.We introduce Cross-Lingual Continued Instruction Tuning (X-CIT), which fully leverages translation-based parallel instruction data to enhance cross-lingual adaptability.X-CIT emulates the human process of second language acquisition and is guided by Chomsky's Principles and Parameters Theory.It first fine-tunes the LLM on English instruction data to establish foundational capabilities (i.e.Principles), then continues with target language translation and customized chatinstruction data to adjust "parameters" specific to the target language.This chat-instruction data captures alignment information in translated parallel data, guiding the model to initially think and respond in its native language before transitioning to the target language.To further mimic human learning progression, we incorporate Self-Paced Learning (SPL) during continued training, allowing the model to advance from simple to complex tasks.Implemented on Llama-2-7B across five languages, X-CIT was evaluated against three objective benchmarks and an LLM-as-a-judge benchmark, improving the strongest baseline by an average of 1.97% and 8.2% in these two benchmarks, respectively.
Linjuan Wu, Baosong Yang, Weiming Lu 0001
ACL (1)1
2025 Enhancing LLM Language Adaption through Cross-lingual In-Context Pre-training
abstract
Large language models (LLMs) exhibit remarkable multilingual capabilities despite Englishdominated pre-training, attributed to crosslingual mechanisms during pre-training.Existing methods for enhancing cross-lingual transfer remain constrained by parallel resources, suffering from limited linguistic and domain coverage.We propose Cross-lingual In-context Pre-training (CrossIC-PT), a simple and scalable approach that enhances cross-lingual transfer by leveraging semantically related bilingual texts via simple next-word prediction.We construct CrossIC-PT samples by interleaving semantic-related bilingual Wikipedia documents into a single context window.To access window size constraints, we implement a systematic segmentation policy to split long bilingual document pairs into chunks while adjusting the sliding window mechanism to preserve contextual coherence.We further extend data availability through a semantic retrieval framework to construct CrossIC-PT samples from web-crawled corpus.Experimental results demonstrate that CrossIC-PT improves multilingual performance on three models (Llama-3.1-8B,Qwen2.5-7B, and Qwen2.5-1.5B)across six target languages, yielding performance gains of 3.79%, 3.99%, and 1.95%, respectively, with additional improvements after data augmentation.
Linjuan Wu, Baosong Yang, Fei Huang 0002, Weiming Lu 0001
EMNLP1
2025 AskToAct: Enhancing LLMs Tool Use via Self-Correcting Clarification
abstract
Xuan Zhang, Yongliang Shen, Zhe Zheng, Linjuan Wu, Wenqi Zhang, Yuchen Yan, Qiuying Peng, Jun Wang, Weiming Lu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Yongliang Shen 0001, Linjuan Wu, Wenqi Zhang 0001, Qiuying Peng, Weiming Lu 0001
EMNLP4
2025 SVGenius: Benchmarking LLMs in SVG Understanding, Editing and Generation
abstract
Large Language Models (LLMs) and Multimodal LLMs have shown promising capabilities for SVG processing, yet existing benchmarks suffer from limited real-world coverage, lack of complexity stratification, and fragmented evaluation paradigms. We introduce SVGenius, a comprehensive benchmark comprising 2,377 queries across three progressive dimensions: understanding, editing, and generation. Built on real-world data from 24 application domains with systematic complexity stratification, SVGenius evaluates models through 8 task categories and 18 metrics. We assess 22 mainstream models spanning different scales, architectures, training paradigms, and accessibility levels. Our analysis reveals that while proprietary models significantly outperform open-source counterparts, all models exhibit systematic performance degradation with increasing complexity, indicating fundamental limitations in current approaches; however, reasoning-enhanced training proves more effective than pure scaling for overcoming these limitations, though style transfer remains the most challenging capability across all model types. SVGenius establishes the first systematic evaluation framework for SVG processing, providing crucial insights for developing more capable vector graphics models and advancing automated graphic design applications. Appendix and supplementary materials (including all data and code) are available at https://zju-real.github.io/SVGenius.
Haolei Xu, Fei Tang 0005, Linjuan Wu, Wenqi Zhang 0001, Guiyang Hou, Yongliang Shen 0001, Weiming Lu 0001, Yueting Zhuang
ACM Multimedia8
2024 Self-Contrast: Better Reflection Through Inconsistent Solving Perspectives
abstract
Wenqi Zhang, Yongliang Shen, Linjuan Wu, Qiuying Peng, Jun Wang, Yueting Zhuang, Weiming Lu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Wenqi Zhang 0001, Yongliang Shen 0001, Linjuan Wu, Qiuying Peng, Yueting Zhuang, Weiming Lu 0001
ACL (1)3
2023 Struct-XLM: A Structure Discovery Multilingual Language Model for Enhancing Cross-lingual Transfer through Reinforcement Learning
abstract
Cross-lingual transfer learning heavily relies on well-aligned cross-lingual representations.The syntactic structure is recognized as beneficial for cross-lingual transfer, but limited researches utilize it for aligning representation in multilingual pre-trained language models (PLMs).Additionally, existing methods require syntactic labels that are difficult to obtain and of poor quality for low-resource languages.To address this gap, we propose Struct-XLM, a novel multilingual language model that leverages reinforcement learning (RL) to autonomously discover universal syntactic structures for improving the cross-lingual representation alignment of PLM.Struct-XLM integrates a policy network (PNet) and a translation ranking task.The PNet is designed to discover structural information and integrate it into the last layer of the PLM through the structural multi-head attention module to obtain structural representation.The translation ranking task obtains a delayed reward based on the structural representation to optimize the PNet while improving the alignment of cross-lingual representation.Experiments show the effectiveness of the proposed approach for enhancing cross-lingual transfer of multilingual PLM on the XTREME benchmark 1 .
Linjuan Wu, Weiming Lu 0001
EMNLP1
2022 Learning Disentangled Semantic Representations for Zero-Shot Cross-Lingual Transfer in Multilingual Machine Reading Comprehension
abstract
Linjuan Wu, Shaojuan Wu, Xiaowang Zhang, Deyi Xiong, Shizhan Chen, Zhiqiang Zhuang, Zhiyong Feng. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Linjuan Wu, Shaojuan Wu, Xiaowang Zhang, Deyi Xiong, Shizhan Chen, Zhiqiang Zhuang, Zhiyong Feng 0002
ACL (1)1
2021 Modeling Global Semantics for Question Answering over Knowledge Bases
abstract
Query graph as a junction of semantic parsing in question answering over knowledge bases (KBQA) connects questions and logical queries. Though query graph consists of rich information such as structure, relation, etc., the current KBQA's models mainly utilize limited relation information in a naive way. It is not easy to learn the representation of a query graph with that information due to the heterogeneity of the query graph and intricate correlation of relations. In this paper, we propose a Global Semantic-based Message Passing (GSMP) model to model the global semantics of a query graph from its structure and relation information. In GSMP, we present a recurrent-based relational graph convolutional network (RGCN) to capture heterogeneous query graphs where the recurrent unit improves the capability of RGCN in processing small-scale query graphs. Moreover, we present a contextual-based method to remove ambiguity caused by intricate correlations where the contextual adjacency of relations optimizes relation representation. Finally, we present a nonlinear gate-based encoder to learning the representation of questions' syntactic tree, as the structure information of questions, for better matching the global semantics of query graphs. Experiments evaluated on benchmarks show that our model outperforms off-the-shelf models.
Peiyun Wu, Yunjie Wu, Linjuan Wu, Xiaowang Zhang, Zhiyong Feng 0002
IJCNN3