Jingqi Tong

dblp:377/9045 · DBLP profile ↗
← Back
7ranked-venue papers
0as first author
7since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 7 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Language models and text generation · 45% Vision and language · 19% Deep learning architectures and training · 16%

Topics — the 18 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
large language model evaluation
1.822026
LLMEval-Fair: A Large-Scale Longitudinal Study on Robust and Fair Evaluation of Large Language Models · ACL (1) 2026
Exploring the Compositional Deficiency of Large Language Models in Mathematical Reasoning Through Trap Problems · EMNLP 2024
Natural language and speech › Language models and text generation › large language model › knowledge in language models
knowledge retention
1.012026
Beyond Scaling: Measuring and Predicting the Upper Bound of Knowledge Retention in Language Model Pre-Training · ACL (1) 2026
Computer vision › Video understanding and tracking
long video understanding
1.012026
VideoPro: Adaptive Program Reasoning for Long Video Understanding · ACL (1) 2026
Computer vision › Vision and language › multimodal reasoning
multi-image reasoning
1.012026
OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Models · ACL (1) 2026
Machine learning › Representation and self-supervised learning
pre-training
1.012026
Beyond Scaling: Measuring and Predicting the Upper Bound of Knowledge Retention in Language Model Pre-Training · ACL (1) 2026
Machine learning › Deep learning architectures and training
scaling laws
1.012026
Beyond Scaling: Measuring and Predicting the Upper Bound of Knowledge Retention in Language Model Pre-Training · ACL (1) 2026
Computer vision › Vision and language
vision-language model
1.012026
OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Models · ACL (1) 2026
Natural language and speech › Language models and text generation › retrieval-augmented generation
knowledge conflict
0.912025
Understanding Parametric and Contextual Knowledge Reconciliation within Large Language Models · NeurIPS 2025
Natural language and speech › Language models and text generation
retrieval-augmented generation
0.912025
Understanding Parametric and Contextual Knowledge Reconciliation within Large Language Models · NeurIPS 2025
Natural language and speech › Language models and text generation
compositional generalization
0.812024
Exploring the Compositional Deficiency of Large Language Models in Mathematical Reasoning Through Trap Problems · EMNLP 2024
Natural language and speech › Language models and text generation › language model analysis
language model scaling
0.812024
LLaMA-MoE: Building Mixture-of-Experts from LLaMA with Continual Pre-Training · EMNLP 2024
Natural language and speech › Language models and text generation
mathematical reasoning
0.812024
Exploring the Compositional Deficiency of Large Language Models in Mathematical Reasoning Through Trap Problems · EMNLP 2024
Machine learning › Deep learning architectures and training
mixture of experts
0.812024
LLaMA-MoE: Building Mixture-of-Experts from LLaMA with Continual Pre-Training · EMNLP 2024
Machine learning › Efficient and distributed learning
model compression
0.812024
LLaMA-MoE: Building Mixture-of-Experts from LLaMA with Continual Pre-Training · EMNLP 2024
Machine learning › Deep learning architectures and training › mixture of experts
sparse expert routing
0.812024
LLaMA-MoE: Building Mixture-of-Experts from LLaMA with Continual Pre-Training · EMNLP 2024
Natural language and speech › Language models and text generation
large language model
0.312026
OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Models · ACL (1) 2026
Machine learning › Trustworthy machine learning
interpretability
0.312025
Understanding Parametric and Contextual Knowledge Reconciliation within Large Language Models · NeurIPS 2025
Knowledge, reasoning and agents › Knowledge representation and reasoning
commonsense reasoning
0.212024
Exploring the Compositional Deficiency of Large Language Models in Mathematical Reasoning Through Trap Problems · EMNLP 2024

Methods — techniques the papers use, named apart from their topics

scaling analysis · 1.0program synthesis · 1.0longitudinal study · 1.0large language model · 1.0benchmark construction · 1.0linear probing · 0.9entity flow · 0.9LoRA · 0.9expert partitioning · 0.8continual pre-training · 0.8
YearPublicationVenuePosition
2026 OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Models
abstract
Qiguang Chen, Chengyu Luan, Jiajun Wu, Qiming Yu, Yi Yang, Yizhuo Li, Jingqi Tong, Xiachong Feng, Libo Qin, Wanxiang Che. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Qiguang Chen, Chengyu Luan, Qiming Yu, Yizhuo Li 0007, Jingqi Tong, Xiachong Feng, Libo Qin 0001, Wanxiang Che
ACL (1)7
2026 Beyond Scaling: Measuring and Predicting the Upper Bound of Knowledge Retention in Language Model Pre-Training
abstract
Changhao Jiang, Ming Zhang, Yifei Cao, Junjie Ye, Xiaoran Fan, Shihan Dou, Zhiheng Xi, Jiajun Sun, Yi Dong, Yujiong Shen, Jingqi Tong, Baoyu Fan, Tao Gui, Qi Zhang, Xuanjing Huang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Changhao Jiang, Ming Zhang 0030, Yifei Cao, Junjie Ye 0005, Xiaoran Fan, Shihan Dou, Zhiheng Xi, Yujiong Shen, Jingqi Tong, Baoyu Fan, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001
ACL (1)11
2026 VideoPro: Adaptive Program Reasoning for Long Video Understanding
abstract
Chenglin Li, Feng Han, Yikun Wang, Ruilin Li, Shuai Dong, Haowen Hou, Haitao Li, Qianglong Chen, Feng Tao, Jingqi Tong, Yin Zhang, Jiaqi Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yikun Wang 0001, Haowen Hou, Qianglong Chen, Jingqi Tong, Yin Zhang 0006
ACL (1)10
2026 LLMEval-Fair: A Large-Scale Longitudinal Study on Robust and Fair Evaluation of Large Language Models
abstract
Ming Zhang, Yujiong Shen, Jingyi Deng, Yuhui Wang, Huayu Sha, Kexin Tan, Qiyuan Peng, Yue Zhang, Junzhe Wang, Shichun Liu, Yueyuan Huang, Jingqi Tong, Changhao Jiang, Yilong Wu, Zhihao Zhang, Mingqi Wu, Mingxu Chai, Zhiheng Xi, Shihan Dou, Tao Gui, Qi Zhang, Xuanjing Huang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Ming Zhang 0030, Yujiong Shen, Jingyi Deng, Huayu Sha, Kexin Tan, Qiyuan Peng, Yue Zhang 0004, Junzhe Wang 0001, Shichun Liu, Yueyuan Huang, Jingqi Tong, Changhao Jiang, Yilong Wu, Zhihao Zhang 0002, Mingqi Wu, Mingxu Chai, Zhiheng Xi, Shihan Dou, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001
ACL (1)12
2025 Understanding Parametric and Contextual Knowledge Reconciliation within Large Language Models
abstract
Retrieval-Augmented Generation (RAG) provides additional contextual knowledge to complement the parametric knowledge in Large Language Models (LLMs). These two knowledge interweave to enhance the accuracy and timeliness of LLM responses. However, the internal mechanisms by which LLMs utilize these knowledge remain unclear. We propose modeling the forward propagation of knowledge as an entity flow, employing this framework to trace LLMs' internal behaviors when processing mixed-source knowledge. Linear probing utilizes a trainable linear classifier to detect specific attributes in hidden layers. However, once trained, a probe cannot adapt to dynamically specified entities. To address this challenge, we construct an entity-aware probe, which introduces special tokens to mark probing targets and employs a small trainable rank-8 lora update to process these special markers. We first verify this approach through an attribution experiment, demonstrating that it can accurately detect information about ad-hoc entities from complex hidden states. Next, we trace entity flows across layers to understand how LLMs reconcile conflicting knowledge internally. Our probing results reveal that contextual and parametric knowledge are routed between tokens through distinct sets of attention heads, supporting attention competition only within knowledge types. While conflicting knowledge maintains a residual presence across layers, aligned knowledge from multiple sources gradually accumulates, with the magnitude of this accumulation directly determining its influence on final outputs.
Jun Zhao 0019, Yongzhuo Yang, Jingqi Tong, Wei Wu 0014, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001
NeurIPS4
2024 LLaMA-MoE: Building Mixture-of-Experts from LLaMA with Continual Pre-Training
abstract
Mixture-of-Experts (MoE) has gained increasing popularity as a promising framework for scaling up large language models (LLMs).However, training MoE from scratch in a largescale setting still suffers from data-hungry and instability problems.Motivated by this limit, we investigate building MoE models from existing dense large language models.Specifically, based on the well-known LLaMA-2 7B model, we obtain an MoE model by: (1) Expert Construction, which partitions the parameters of original Feed-Forward Networks (FFNs) into multiple experts; (2) Continual pretraining, which further trains the transformed MoE model and additional gate networks.In this paper, we comprehensively explore different methods for expert construction and various data sampling strategies for continual pretraining.After these stages, our LLaMA-MoE models could maintain language abilities and route the input tokens to specific experts with part of the parameters activated.Empirically, by training 200B tokens, LLaMA-MoE-3.5Bmodels significantly outperform dense models that contain similar activation parameters.
Tong Zhu 0002, Xiaoye Qu, Daize Dong, Jiacheng Ruan, Jingqi Tong, Conghui He, Yu Cheng 0001
EMNLP5
2024 Exploring the Compositional Deficiency of Large Language Models in Mathematical Reasoning Through Trap Problems
abstract
Human cognition exhibits systematic compositionality, the algebraic ability to generate infinite novel combinations from finite learned components, which is the key to understanding and reasoning about complex logic.In this work, we investigate the compositionality of large language models (LLMs) in mathematical reasoning.Specifically, we construct a new dataset MATHTRAP ‡ by introducing carefully designed logical traps into the problem descriptions of MATH and GSM8K.Since problems with logical flaws are quite rare in the real world, these represent "unseen" cases to LLMs.Solving these requires the models to systematically compose (1) the mathematical knowledge involved in the original problems with (2) knowledge related to the introduced traps.Our experiments show that while LLMs possess both components of requisite knowledge, they do not spontaneously combine them to handle these novel cases.We explore several methods to mitigate this deficiency, such as natural language prompts, few-shot demonstrations, and fine-tuning.Additionally, we test the recently released OpenAI o1 model and find that human-like 'slow thinking' helps improve the compositionality of LLMs.Overall, systematic compositionality remains an open challenge for large language models.
Jun Zhao 0019, Jingqi Tong, Yurong Mou, Ming Zhang 0030, Qi Zhang 0001, Xuanjing Huang 0001
EMNLP2