Tongxu Luo

dblp:356/7501 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
4since 2021 · last 2026
0009-0000-4576-1178ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Efficient and distributed learning · 42% Vision and language · 21% Language models and text generation · 18%

Topics — the 14 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
inference efficiency
1.012026
Cut Your Losses! Learning to Prune Paths Early for Efficient Parallel Reasoning · ACL (1) 2026
Natural language and speech › Language models and text generation › large language model
large reasoning model
1.012026
Cut Your Losses! Learning to Prune Paths Early for Efficient Parallel Reasoning · ACL (1) 2026
Machine learning › Efficient and distributed learning › efficient training
efficient pre-training
0.812024
Stacking Your Transformers: A Closer Look at Model Growth for Efficient LLM Pre-Training · NeurIPS 2024
Natural language and speech › Language models and text generation › large language model training › language model pretraining
large language model pretraining
0.812024
Stacking Your Transformers: A Closer Look at Model Growth for Efficient LLM Pre-Training · NeurIPS 2024
Machine learning › Efficient and distributed learning › parameter-efficient fine-tuning
low-rank adaptation
0.812024
Neeko: Leveraging Dynamic LoRA for Efficient Multi-Character Role-Playing Agent · EMNLP 2024
Machine learning › Efficient and distributed learning › dynamic neural network
model growth
0.812024
Stacking Your Transformers: A Closer Look at Model Growth for Efficient LLM Pre-Training · NeurIPS 2024
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning
0.812024
Neeko: Leveraging Dynamic LoRA for Efficient Multi-Character Role-Playing Agent · EMNLP 2024
Natural language and speech › Question answering and dialogue systems › open-domain dialogue
role-playing dialogue agents
0.812024
Neeko: Leveraging Dynamic LoRA for Efficient Multi-Character Role-Playing Agent · EMNLP 2024
Natural language and speech › Question answering and dialogue systems
multi-hop reasoning
0.712023
Answer-Based Entity Extraction and Alignment for Visual Text Question Answering · ACM Multimedia 2023
Computer vision › Vision and language
multimodal reasoning
0.712023
Answer-Based Entity Extraction and Alignment for Visual Text Question Answering · ACM Multimedia 2023
Computer vision › Vision and language › visual question answering
text-based visual question answering
0.712023
Answer-Based Entity Extraction and Alignment for Visual Text Question Answering · ACM Multimedia 2023
Computer vision › Vision and language
visual question answering
0.712023
Answer-Based Entity Extraction and Alignment for Visual Text Question Answering · ACM Multimedia 2023
Machine learning › Deep learning architectures and training
transformer
0.212024
Stacking Your Transformers: A Closer Look at Model Growth for Efficient LLM Pre-Training · NeurIPS 2024
Natural language and speech › Information extraction and text analysis
named entity recognition
0.212023
Answer-Based Entity Extraction and Alignment for Visual Text Question Answering · ACM Multimedia 2023

Methods — techniques the papers use, named apart from their topics

token-level pruning · 1.0reinforcement learning · 1.0model growth operators · 0.8gating network · 0.8dynamic LoRA · 0.8depthwise stacking · 0.8entity extraction · 0.7cross-modal alignment · 0.7
YearPublicationVenuePosition
2026 Cut Your Losses! Learning to Prune Paths Early for Efficient Parallel Reasoning
abstract
Parallel reasoning enhances Large Reasoning Models (LRMs) but incurs prohibitive costs due to futile paths caused by early errors.To mitigate this, path pruning at the prefix level is essential, yet existing research remains fragmented without a standardized framework.In this work, we propose the first systematic taxonomy of path pruning, categorizing methods by their signal source (internal vs. external) and learnability (learnable vs. non-learnable).This classification reveals the unexplored potential of learnable internal methods, motivating our proposal of STOP (Super TOken for Pruning).Extensive evaluations across LRMs ranging from 1.5B to 20B parameters demonstrate that STOP achieves superior effectiveness and efficiency compared to existing baselines.Furthermore, we rigorously validate the scalability of STOP under varying compute budgets-for instance, boosting GPT-OSS-20B accuracy on AIME25 from 84% to nearly 90% under fixed compute budgets.Finally, we distill our findings into formalized empirical guidelines to facilitate optimal real-world deployment.Code, data and models are available at https://bijiaxihh.github.io/STOP.
Jiaxi Bi, Tongxu Luo, Wenyu Du, Zhengyang Tang, Benyou Wang
ACL (1)2
2024 Neeko: Leveraging Dynamic LoRA for Efficient Multi-Character Role-Playing Agent
abstract
Large Language Models (LLMs) have revolutionized open-domain dialogue agents but encounter challenges in multi-character roleplaying (MCRP) scenarios.To address this issue, this work presents Neeko, an innovative framework designed for efficient multiplecharacter role-playing.The proposed framework breaks down the role-playing agent's training process into agent pre-tuning, multiple character playing, and character incremental learning, effectively handling both seen and unseen roles.Neeko employs a dynamic low-rank adapter (LoRA) strategy by training separate LoRA blocks independently for each character, alongside incorporating a gating network for role selection.This design allows Neeko to seamlessly adjust to a wide range of characters, thereby bolstering its adaptability to distinctive attributes, personalities, and speech patterns.As a result, Neeko demonstrates superior performance in MCRP over most existing methods, offering more engaging and versatile user interaction experiences.
Tongxu Luo, Yifan Wei 0001, Fangyu Lei, Hao Peng 0001, Liehuang Zhu
EMNLP2
2024 Stacking Your Transformers: A Closer Look at Model Growth for Efficient LLM Pre-Training
abstract
LLMs are computationally expensive to pre-train due to their large scale. Model growth emerges as a promising approach by leveraging smaller models to accelerate the training of larger ones. However, the viability of these model growth methods in efficient LLM pre-training remains underexplored. This work identifies three critical $\underline{\textit{O}}$bstacles: ($\textit{O}$1) lack of comprehensive evaluation, ($\textit{O}$2) untested viability for scaling, and ($\textit{O}$3) lack of empirical guidelines. To tackle $\textit{O}$1, we summarize existing approaches into four atomic growth operators and systematically evaluate them in a standardized LLM pre-training setting. Our findings reveal that a depthwise stacking operator, called $G_{\text{stack}}$, exhibits remarkable acceleration in training, leading to decreased loss and improved overall performance on eight standard NLP benchmarks compared to strong baselines. Motivated by these promising results, we conduct extensive experiments to delve deeper into $G_{\text{stack}}$ to address $\textit{O}$2 and $\textit{O}$3. For $\textit{O}$2 (untested scalability), our study shows that $G_{\text{stack}}$ is scalable and consistently performs well, with experiments up to 7B LLMs after growth and pre-training LLMs with 750B tokens. For example, compared to a conventionally trained 7B model using 300B tokens, our $G_{\text{stack}}$ model converges to the same loss with 194B tokens, resulting in a 54.6\% speedup. We further address $\textit{O}$3 (lack of empirical guidelines) by formalizing guidelines to determine growth timing and growth factor for $G_{\text{stack}}$, making it practical in general LLM pre-training. We also provide in-depth discussions and comprehensive ablation studies of $G_{\text{stack}}$. Our code and pre-trained model are available at https://llm-stacking.github.io/.
Wenyu Du, Tongxu Luo, Zihan Qiu, Yikang Shen, Reynold Cheng, Yike Guo, Jie Fu 0001
NeurIPS2
2023 Answer-Based Entity Extraction and Alignment for Visual Text Question Answering
abstract
As a variant of visual question answering (VQA), visual text question answering (VTQA) provides a text-image pair for each question. Text utilizes named entities to describe corresponding image. Consequently, the ability to perform multi-hop reasoning using named entities between text and image becomes critically important. However, existing models pay relatively less attention to this aspect. Therefore, we propose Answer-Based Entity Extraction and Alignment Model (AEEA) to enable a comprehensive understanding and support multi-hop reasoning. The core of AEEA lies in two main components: AKECMR and answer aware predictor. The former emphasizes the alignment of modalities and effectively distinguishes between intra-modal and inter-modal information, and the latter prioritizes the full utilization of intrinsic semantic information contained in answers during training. Our model outperforms the baseline by 2.24% on test-dev set and 1.06% on test set, securing the third place in VTQA2023(English).
Jun Yu 0001, Mohan Jing, Weihao Liu 0004, Tongxu Luo, Keda Lu, Fangyu Lei, Jianqing Sun, Jiaen Liang
ACM Multimedia4